his NSF-funded project studies persuasive AI and the spread of misinformation. One thread of it is a classifier that scores text across four rhetorical strategies, causal, empirical, emotional, and moral, fine-tuned on LLM-labeled synthetic debates and validated against human annotators. The underlying classifier framework has been published at NAACL and EMNLP 2025.
What I did
- Wrote a GDELT scraper from scratch with concurrent pipeline jobs, retry logic, resumable progress tracking, and content validation to maximize clean-text yield.
- Built the corpus: roughly 2 million articles across 14 U.S. outlets spanning 2000 to 2025, cleaned, deduplicated, and standardized into a unified schema of outlet lean, credibility, year, title, and content.
- The corpus became the substrate for the classifier's large-scale media analysis, articles chunked, scored across all four strategies, and aggregated by outlet, year, and genre.
Findings
Low-credibility outlets rely substantially more on moral and emotional appeals than high-credibility outlets, which lean toward causal reasoning. Ideological differences emerge primarily along the same affective dimension. Within mainstream opinion journalism, The New York Times leans more on causal, empirical, and moral argumentation while the New York Post leans more on emotional appeals, with partial convergence between the two over time.