his NSF-funded project studies persuasive AI and the spread of misinformation. One thread of it is a classifier that scores text across four rhetorical strategies, causal, empirical, emotional, and moral, fine-tuned on LLM-labeled synthetic debates and validated against human annotators. The underlying classifier framework has been published at NAACL and EMNLP 2025.

What I did

  • Wrote a GDELT scraper from scratch with concurrent pipeline jobs, retry logic, resumable progress tracking, and content validation to maximize clean-text yield.
  • Built the corpus: roughly 2 million articles across 14 U.S. outlets spanning 2000 to 2025, cleaned, deduplicated, and standardized into a unified schema of outlet lean, credibility, year, title, and content.
  • The corpus became the substrate for the classifier's large-scale media analysis, articles chunked, scored across all four strategies, and aggregated by outlet, year, and genre.
Terminal showing the GDELT scraper starting a run for 2022 with the political-only filter, downloading export files
°01 The scraper kicking off a run: GDELT exports downloading, filtered to political news
Terminal showing filtered URLs written to CSV and the article scraping stage running with OK confirmations
°02 Filtered URLs written to CSV, then article scraping in progress

Findings

Low-credibility outlets rely substantially more on moral and emotional appeals than high-credibility outlets, which lean toward causal reasoning. Ideological differences emerge primarily along the same affective dimension. Within mainstream opinion journalism, The New York Times leans more on causal, empirical, and moral argumentation while the New York Post leans more on emotional appeals, with partial convergence between the two over time.