118. Angela Fan - Generating Wikipedia articles with AI

118. Angela Fan - Generating Wikipedia articles with AI

Author: The TDS team April 6, 2022 Duration: 51:44

Generating well-referenced and accurate Wikipedia articles has always been an important problem: Wikipedia has essentially become the Internet's encyclopedia of record, and hundreds of millions of people use it do understand the world.

But over the last decade Wikipedia has also become a critical source of training data for data-hungry text generation models. As a result, any shortcomings in Wikipedia’s content are at risk of being amplified by the text generation tools of the future. If one type of topic or person is chronically under-represented in Wikipedia’s corpus, we can expect generative text models to mirror — or even amplify — that under-representation in their outputs.

Through that lens, the project of Wikipedia article generation is about much more than it seems — it’s quite literally about setting the scene for the language generation systems of the future, and empowering humans to guide those systems in more robust ways.

That’s why I wanted to talk to Meta AI researcher Angela Fan, whose latest project is focused on generating reliable, accurate, and structured Wikipedia articles. She joined me to talk about her work, the implications of high-quality long-form text generation, and the future of human/AI collaboration on this episode of the TDS podcast.

--- 

Intro music:

- Artist: Ron Gelinas

- Track Title: Daybreak Chill Blend (original mix)

- Link to Track: https://youtu.be/d8Y2sKIgFWc

---

Chapters:

  • 1:45 Journey into Meta AI
  • 5:45 Transition to Wikipedia
  • 11:30 How articles are generated
  • 18:00 Quality of text
  • 21:30 Accuracy metrics
  • 25:30 Risk of hallucinated facts
  • 30:45 Keeping up with changes
  • 36:15 UI/UX problems
  • 45:00 Technical cause of gender imbalance
  • 51:00 Wrap-up

While the active production of Towards Data Science has concluded, its archive remains a vital resource. Created by The TDS team, this collection captures a specific moment in the rapid evolution of data science and artificial intelligence. Each conversation pulls you directly into the room with leading researchers and practitioners who were shaping the tools and theories of their time. The discussions are not abstract lectures; they are grounded explorations of real-world problems, ethical dilemmas, and technical challenges that defined the field's trajectory. You'll hear experts dissect the implications of their work, from algorithmic fairness to the practicalities of deploying models at scale. This podcast served as a forum for nuanced debate, where complex ideas were unpacked with clarity and depth. Listening now offers a unique historical perspective, a chance to understand the foundational conversations that continue to influence where technology is headed next. The archive of Towards Data Science stands as a substantive record of insight, preserving the voices and questions from the forefront of a digital revolution.
Author: Language: en-us Episodes: 50

Towards Data Science
Podcast Episodes
119. Jaime Sevilla - Projecting AI progress from compute trends [not-audio_url] [/not-audio_url]

Duration: 48:34
There’s an idea in machine learning that most of the progress we see in AI doesn’t come from new algorithms of model architectures. instead, some argue, progress almost entirely comes from scaling up compute power, datas…
117. Beena Ammanath - Defining trustworthy AI [not-audio_url] [/not-audio_url]

Duration: 46:46
Trustworthy AI is one of today’s most popular buzzwords. But although everyone seems to agree that we want AI to be trustworthy, definitions of trustworthiness are often fuzzy or inadequate. Maybe that shouldn’t be surpr…
116. Katya Sedova - AI-powered disinformation, present and future [not-audio_url] [/not-audio_url]

Duration: 54:24
Until recently, very few people were paying attention to the potential malicious applications of AI. And that made some sense: in an era where AIs were narrow and had to be purpose-built for every application, you’d need…
115. Irina Rish - Out-of-distribution generalization [not-audio_url] [/not-audio_url]

Duration: 50:12
Imagine, for example, an AI that’s trained to identify cows in images. Ideally, we’d want it to learn to detect cows based on their shape and colour. But what if the cow pictures we put in the training dataset always sho…
114. Sam Bowman - Are we *under-hyping* AI? [not-audio_url] [/not-audio_url]

Duration: 47:48
Google the phrase “AI over-hyped”, and you’ll find literally dozens of articles from the likes of Forbes, Wired, and Scientific American, all arguing that “AI isn’t really as impressive at it seems from the outside,” and…
113. Yaron Singer - Catching edge cases in AI [not-audio_url] [/not-audio_url]

Duration: 35:20
It’s no secret that AI systems are being used in more and more high-stakes applications. As AI eats the world, it’s becoming critical to ensure that AI systems behave robustly — that they don’t get thrown off by unusual…
110. Alex Turner - Will powerful AIs tend to seek power? [not-audio_url] [/not-audio_url]

Duration: 46:57
Today’s episode is somewhat special, because we’re going to be talking about what might be the first solid quantitative study of the power-seeking tendencies that we can expect advanced AI systems to have in the future.…