Research
PhD researcher building production LLM pipelines for computational social science.
Expertise in HPC deployment, evaluation methodology, and causal inference, with a record
of reducing weeks of manual research labor to hours of parallelized computation and
extending causal estimators to cases standard methods cannot handle.
LLM Methodology for Social Science
Under ReviewLLM AnnotationHPCNSF FundedText-as-DataData Annotation
Using Messy Text in Future LLM Annotations
Researchers often inherit annotations built over years of costly human coding, but noise
and ambiguity in older corpora make them unusable for modern text-as-data pipelines.
How can that accumulated knowledge be reused without starting over? This research is supported by an NSF grant.
- Deployed an LLM pipeline on HPC that uses codebook-guided attention to realign annotations with noisy text, producing clean, evidence-traceable summaries for downstream tasks.
- Enables researchers to reuse annotations acquired over years of costly human coding, reducing weeks of manual labor to hours of parallelized computation without requiring re-annotation.
Under ReviewHyperspectral ImagingArcGISMultimodal LLMHumanitarian DeminingRemote SensingComputer VisionBenchmarking
UAV Hyperspectral Landmine and UXO Detection: ArcGIS vs Multimodal LLM
Can AI systems match classical geospatial methods for detecting landmines in aerial imagery?
The paper runs a direct head-to-head comparison on a humanitarian demining field survey.
- Benchmarks ArcGIS hyperspectral anomaly detection against a locally-served multimodal LLM on UAV aerial imagery from a humanitarian mine and UXO field survey.
- Establishes an empirical comparison between classical geospatial ML and AI-assisted detection on a safety-critical task where false negatives have lethal consequences.
Working PaperLLM EvaluationBenchmarkingHPC
Biases Evaluation and Scope of LLM Applications in Social Science Corpora
Existing LLM benchmarks use QA and multiple-choice formats on low-risk texts. Social
scientists work with annotation, extraction, classification, and simulation on corpora
where risk levels vary, from routine survey data to conflict, crime, and political
violence. This project measures how much LLM bias degrades research results across
tasks and risk levels.
- Benchmarks LLM bias magnitude across research tasks and corpora with varying risk levels, not the QA and multiple-choice formats that dominate existing benchmarks.
- Produces scope lines identifying which task-text pairs tolerate unmitigated LLM bias for research use, replacing trial-and-error with reproducible, task-specific thresholds.
Presented at ConferenceInformation ExtractionNamed Entity ExtractionComparative PoliticsBenchmarking
LLM-Based Rhetoric Extraction from Censorship-Related Legal Texts
What language do governments use in censorship laws, and does it vary systematically with
how tightly they control internet access and expression? This paper builds pipelines to extract rights- and
security-related rhetoric from censorship legislation and measure it at cross-national scale.
- Benchmarks two extraction pipelines against each other, establishing a reproducible, auditable method for measuring censorship rhetoric across national legal corpora.
- Span-verifiable annotations make every measurement checkable and the cross-national comparison reproducible at scale.
Under ReviewAI PolicyState LegislationAgentic RAGFine-TuningInformation RetrievalNamed Entity RecognitionAI Governance
Predicting AI Policy in U.S. State Legislation
What predicts how U.S. states define and regulate AI in their legislation? This project
deploys an agentic RAG pipeline to extract regulated entities from state AI bills,
fine-tunes a model to predict policy adoption, and benchmarks model bias on the corpus.
- An agentic RAG pipeline extracts regulated entities from state AI legislation; a fine-tuned model predicts policy adoption and quantifies model bias on the corpus.
- The pipeline generalizes across corpora and domains, making bias-quantified AI-assisted policy analysis scalable without redesign.
Causal InferenceMedia BiasDemocratic BackslidingPolitical CommunicationAgenda Setting
Media Bias and Democratic Backsliding
Do political leaders use rhetoric to distort media narratives, and can this enable democratic backsliding? This paper estimates whether presidential social media posts about crime causally shift regional news coverage, treating elite rhetoric as an intervention in a causal design.
- Provides causal evidence that presidential social media rhetoric shifts regional news coverage, widening the gap between actual crime rates and media representation.
- Connects media bias to democratic backsliding: distorted information environments reduce perceived government performance, creating conditions under which voters support disruptive political change.
Causal Inference & Methodology
Presented at ConferenceSynthetic ControlExport ControlsAI Competition
Contextualized Synthetic Control: The NVIDIA Chip Ban and the AI Arms Race
Standard synthetic control requires the treated unit to be comparable to others in the
donor pool. When that condition fails, as it does for China in the AI arms race, the method
cannot produce valid counterfactuals. This paper extends the estimator to handle cases
where no comparable donors exist.
- Augments synthetic control with multi-head attention to produce valid counterfactuals when the treated unit has no comparable donors.
- Relaxes the donor pool comparability requirement, extending causal inference to cases previously excluded by scale disparity.
Presented at ConferenceUnder ReviewDifference-in-DifferencesExport ControlsAI CompetitionExport Regulation
Hardware Restrictions and AI Competition: The Effect of NVIDIA GPU Export Restrictions on AI Benchmark Performance
Did U.S. export controls on high-end GPUs actually slow China's AI development? The paper
estimates the effect of successive export control waves on Chinese AI model performance
over time.
- Builds a developer-level monthly panel from the Hugging Face Open LLM Leaderboard and estimates the effect of successive export control waves with staggered difference-in-differences and mechanism-specific extensions.
- Finds no robust negative net effect on Chinese benchmark performance across the main estimators; mechanism tests are more consistent with adaptation (via substitute GPUs, organizational scale, and increased research output) than with suppression.