AI Safety

New ML pipeline turns Twitter data into public policy agendas

Researchers combine GPT-2, LDA, and human validation to mine public issues from social media.

Deep Dive

A new research paper from Rahman Sanya introduces a human-augmented machine learning framework for automatically building public policy agendas from social media conversations. The approach aims to replace traditional time- and labor-intensive methods of issue identification and agenda setting with a scalable five-stage pipeline: input cleaning, keyword extraction and issue identification, narrative creation, narrative validation, and agenda validation.

The experiments were conducted on a Twitter dataset using Latent Dirichlet Allocation (LDA) and Top2Vec for topic modeling, GPT-2 for natural language generation (NLG), and both similarity analysis and human evaluation for validation. Results showed 'very good' and 'good' inter-rater agreement (IRA) for readability and coherence of generated narratives, and 'good' IRA for agenda items. Cosine similarity scores were above average for at least three out of five reference themes, demonstrating the promise of ML for efficient, large-scale public policy agenda setting.

Key Points
  • Five-stage ML pipeline: data cleaning, keyword extraction, narrative generation (GPT-2), narrative validation, agenda validation.
  • Uses LDA and Top2Vec for topic modeling on Twitter data to surface public issues.
  • Achieved 'very good' inter-rater agreement on narrative readability and coherence; above-average cosine similarity on 3 of 5 themes.

Why It Matters

Scalable, automated issue identification from social media could transform how governments and organizations set policy priorities.

📬 Get the top 10 AI stories daily