Research

PhD researcher building production LLM pipelines for computational social science. Expertise in HPC deployment, evaluation methodology, and causal inference, with a record of reducing weeks of manual research labor to hours of parallelized computation and extending causal estimators to cases standard methods cannot handle.

LLM Methodology for Social Science

Under ReviewLLM AnnotationHPCNSF FundedText-as-DataData Annotation

Using Messy Text in Future LLM Annotations

Researchers often inherit annotations built over years of costly human coding, but noise and ambiguity in older corpora make them unusable for modern text-as-data pipelines. How can that accumulated knowledge be reused without starting over? This research is supported by an NSF grant.

  • Deployed an LLM pipeline on HPC that uses codebook-guided attention to realign annotations with noisy text, producing clean, evidence-traceable summaries for downstream tasks.
  • Enables researchers to reuse annotations acquired over years of costly human coding, reducing weeks of manual labor to hours of parallelized computation without requiring re-annotation.
Under ReviewHyperspectral ImagingArcGISMultimodal LLMHumanitarian DeminingRemote SensingComputer VisionBenchmarking

UAV Hyperspectral Landmine and UXO Detection: ArcGIS vs Multimodal LLM

Can AI systems match classical geospatial methods for detecting landmines in aerial imagery? The paper runs a direct head-to-head comparison on a humanitarian demining field survey.

  • Benchmarks ArcGIS hyperspectral anomaly detection against a locally-served multimodal LLM on UAV aerial imagery from a humanitarian mine and UXO field survey.
  • Establishes an empirical comparison between classical geospatial ML and AI-assisted detection on a safety-critical task where false negatives have lethal consequences.
Working PaperLLM EvaluationBenchmarkingHPC

Biases Evaluation and Scope of LLM Applications in Social Science Corpora

Existing LLM benchmarks use QA and multiple-choice formats on low-risk texts. Social scientists work with annotation, extraction, classification, and simulation on corpora where risk levels vary, from routine survey data to conflict, crime, and political violence. This project measures how much LLM bias degrades research results across tasks and risk levels.

  • Benchmarks LLM bias magnitude across research tasks and corpora with varying risk levels, not the QA and multiple-choice formats that dominate existing benchmarks.
  • Produces scope lines identifying which task-text pairs tolerate unmitigated LLM bias for research use, replacing trial-and-error with reproducible, task-specific thresholds.
Presented at ConferenceInformation ExtractionNamed Entity ExtractionComparative PoliticsBenchmarking

LLM-Based Rhetoric Extraction from Censorship-Related Legal Texts

What language do governments use in censorship laws, and does it vary systematically with how tightly they control internet access and expression? This paper builds pipelines to extract rights- and security-related rhetoric from censorship legislation and measure it at cross-national scale.

  • Benchmarks two extraction pipelines against each other, establishing a reproducible, auditable method for measuring censorship rhetoric across national legal corpora.
  • Span-verifiable annotations make every measurement checkable and the cross-national comparison reproducible at scale.
Under ReviewAI PolicyState LegislationAgentic RAGFine-TuningInformation RetrievalNamed Entity RecognitionAI Governance

Predicting AI Policy in U.S. State Legislation

What predicts how U.S. states define and regulate AI in their legislation? This project deploys an agentic RAG pipeline to extract regulated entities from state AI bills, fine-tunes a model to predict policy adoption, and benchmarks model bias on the corpus.

  • An agentic RAG pipeline extracts regulated entities from state AI legislation; a fine-tuned model predicts policy adoption and quantifies model bias on the corpus.
  • The pipeline generalizes across corpora and domains, making bias-quantified AI-assisted policy analysis scalable without redesign.
Causal InferenceMedia BiasDemocratic BackslidingPolitical CommunicationAgenda Setting

Media Bias and Democratic Backsliding

Do political leaders use rhetoric to distort media narratives, and can this enable democratic backsliding? This paper estimates whether presidential social media posts about crime causally shift regional news coverage, treating elite rhetoric as an intervention in a causal design.

  • Provides causal evidence that presidential social media rhetoric shifts regional news coverage, widening the gap between actual crime rates and media representation.
  • Connects media bias to democratic backsliding: distorted information environments reduce perceived government performance, creating conditions under which voters support disruptive political change.

Causal Inference & Methodology

Presented at ConferenceSynthetic ControlExport ControlsAI Competition

Contextualized Synthetic Control: The NVIDIA Chip Ban and the AI Arms Race

Standard synthetic control requires the treated unit to be comparable to others in the donor pool. When that condition fails, as it does for China in the AI arms race, the method cannot produce valid counterfactuals. This paper extends the estimator to handle cases where no comparable donors exist.

  • Augments synthetic control with multi-head attention to produce valid counterfactuals when the treated unit has no comparable donors.
  • Relaxes the donor pool comparability requirement, extending causal inference to cases previously excluded by scale disparity.
Presented at ConferenceUnder ReviewDifference-in-DifferencesExport ControlsAI CompetitionExport Regulation

Hardware Restrictions and AI Competition: The Effect of NVIDIA GPU Export Restrictions on AI Benchmark Performance

Did U.S. export controls on high-end GPUs actually slow China's AI development? The paper estimates the effect of successive export control waves on Chinese AI model performance over time.

  • Builds a developer-level monthly panel from the Hugging Face Open LLM Leaderboard and estimates the effect of successive export control waves with staggered difference-in-differences and mechanism-specific extensions.
  • Finds no robust negative net effect on Chinese benchmark performance across the main estimators; mechanism tests are more consistent with adaptation (via substitute GPUs, organizational scale, and increased research output) than with suppression.