The following is a write up of a keynote I was invited to give at SciCAR, the annual German gathering of scientists, journalists, and researchers. Consider it under-construction thoughts, aligned with the tradition of blogging to think out loud.
Everyday we see a new gaffe, book, lawsuit, or report about AI and journalism. In the practice of data journalism, like many others, this new technology is already influencing how we approach learning in newsrooms and classroom. I want to pause to reflect my own approaches and catalog some inspirational examples here to tease out how we can ensure any engagement with AI in data journalism work is in line with the fundamental principles that underlie the practice of being a journalist.
SciCAR, like NICAR and the Computation +Journalism Symposium, is one of the few places where data journalists, academics, scientists and others come together to discuss and develop our collaborative work. You’ll note that NICAR and SciCAR share a now-antiquated acronym: “Computer Assisted Reporting” (CAR). The term sounds a bit dated to the ear because there few areas of reporting that aren’t assisted by computers, but “CAR” reminds us that computing was once considered a radical and unfamiliar force in newsrooms. CAR came to popularity as a label for exploring the possibilities of how computers could help in reporting, and the ways it interfered with the regular practice of journalism.
I was honored to be invited to deliver this keynote talk at SciCAR in Dortumnd.
I think we need to return to that way of thinking in an era of chat-based large language models (LLMs), considering how and what we want computers to assist with. So CAR is the basis of my somewhat tongue-in-cheek reminder: we must teach and learn extra CAR-efully right now, committed to thoughtful consideration of how and where we use AI to assist our data reporting. See what I did there?
Some Context
Like so many of you I’ve been tracking developments in AI and journalism, helping launch a series of related news round-ups on Storybench that our Northeastern J-School grad student Vivica D’Souza continues to release every month or two. A core tension emerged for me from watching developments over the last two years: practicing journalists are both adapting and resisting. In parallel, universities that train new journalists are widely adopting AI, causing many student groups to organize and push back (arguing that “AI is adverse to the idea of learning”).
The flood of news updates about AI and journalism is almost impossible to keep up with. We’ve been writing round-ups on Storybench to highlight developments every month or two.
I interpret this resistance as centrally about power: who gets to decide how and for what these new tools are used in newsrooms and J-School classrooms? With this lens I’m immediately reminded of the Luddites of the early 1800s. Luddites weren’t opposed to the idea of industrialization, they were skilled workers fighting for the application of those new machines in ways that continued to produce high quality goods, were under the control of workers, and didn’t violate existing labor practices.
Does that list sound familiar? It sure does to me. Luddites come up often in media discussions about resistance to AI. In the field of data journalism, AI-driven work risks polluting product quality, is being forced into many learning settings, and is putting careers at risk. Yet practicing data journalists are rapidly starting to use AI to help develop stories, create interactive visuals, support investigative analysis, and more. There’s that central tension again.
So I choose to engage with AI technologies in my teaching and collaborations with learners of all types. I focus on thoughtful use of AI, critiquing misuse of AI, and principled refuse of AI. Data journalists have agency in how AI is engaged, or not engaged with, in their work. Let me run through a few central principles and help me reflect on how to do this well.
Be Specific In Your Learning
Once of my guiding principles here is to always be specific when I introduce AI for data journalism. AI companies market that their models can and will help with everything, but that’s mostly because they are unproven and untested; these are new tools that are looking for market validation and dominance. In that way they remind me of the introduction of the Apple Watch, which was pitched as a communication device (remember the idea of drawing hearts to send to your loved one?), a fashion accessory (remember the 24 karat gold watch?), a high-precision tool to tell time and a health and fitness tracker. A decade later it is clear that the last on those list is the market they found; it is primarily a health and fitness wearable.
AI hasn’t gotten there yet. We don’t really know yet which applications it will really help us with. This makes it hard to introduce in productive ways. Can AI help us brainstorm, automate workflows, edit our work, code for us, scrape datasets, make charts, add annotations, surface insights in CSVs? We don’t know yet. I find that being specific about those as separate potential capabilities helps significantly.
We’re beginning to see examples of recurring concrete use-cases for AI in data journalism, which I classify into buckets of data production, data shaping, data analysis, and data representation. One example on the data shaping side is CNN’s recent work identifying a pattern of Trump promoting companies on his Truth Social blog days after buying their stocks. Reporter Isabelle Chapman shared how they “used artificial intelligence to compare a database of Trump’s Truth Social posts with a full list of stock trades released in his annual financial disclosure”. I imagine using AI to link between the two datasets helped them with scale and fuzzy matching logic. These two tasks are well-known challenges in investigative data work, so they offer very specific (and verifiable) opportunities to leverage AI.
Another example, from the realm of data representation, is using AI to write the code that renders complex and novel interactive graphics. This past spring graphics developer (and Northeastern J-School grad) Yuriko Schumacher shared with my Data Storytelling class how she’s using AI coding in her work at the Athletic. Code generation has helped their teams iterate faster, robustly support desktop and mobile, and focus on interactive design so their readers are informed and engaged with complex visualizations that represent large sports-related datasets.
I point to these kinds of examples in my teaching to support discussion about how AI might help or hurt our ability to perform specific parts of our data storytelling. We don’t discuss “AI” as an abstract; we instead focus on concrete applications of chat-based LLMs within a list of concrete data journalism tasks.
Learn by Trying Things Out
Specific examples help push against the pervasive industry hype about AI “changing everything,” offering a path to concrete things learners can try out. Scholars from various domains have led research into how organizations working for the public good often repurpose technologies, putting them into service for purposes other than those they were designed for (including in journalism). Over and over this research shows that the process of adapting technological tools begins by simply trying something out to see if it helps.
One example from my own teaching is related to infographics. Early in my Data Storytelling course I ask students to write a short story about a small dataset with two standard charts. After this I have them put a one paragraph narrative summary, alongside the dataset, into a nano-banana model (via Gemini), asking it to turn their story into an infographic. The results are mixed and often hilarious, prompting opportunities for visual critique and sometimes sparks of symbols that could be repurposed. Most of the graphics produced last year were terrible but the point is that we practiced assessing some hype: nano-banana is supposed to be great at infographics, but the results we got by trying it out quickly were unimpressive. That’s important, and builds confidence in students to try things out and make their own judgments rather than trusting the marketing hype.
Examples of infographics created by my students using Gemini in spring of 2026.
Trying things out like this sounds great, but we have to consider the costs incurred in the playgrounds we create for ourselves and others. Legitimate and real concerns about water use, energy supply costs, training data copyright theft, big tech motivations, increasing costs, the spread of “slop” AI generated content, out-of-date models, and impact on cognition are all raised regularly and written about in depth. These can’t be ignored and have to continually inform how much we’re willing to try out AI tools that have so many negative impacts. I don’t have an easy answer here, so instead I engage with my peers and students to see how they feel vis-a-vis these concerns after trying out an AI tool.
Shaw & Nave in their oft-linked “cognitive surrender” paper found that people using AI to assist a task just kind of did what it told them.
However, I do think trying to generate code is a distinctly specific task within that space. Yes, I know and have seen the mixed results from research on learning to code with AI, but put that aside for a minute. My concerns about those costs I listed above are reduced in the case of AI for coding for a few reasons. First, much of the training material was obtained legally, scraped from online open source repositories with licenses. Second, data representation or analysis in journalism contexts is short-lived, with a ship-date after which code seldom needs to be revisited so technical debt accrued doesn’t matter that much. Third, local models are getting better quickly, meaning that cloud costs and corporate control aren’t as relevant. Finally, code is verifiable. You can see what the AI model produced, and the code you run is deterministic, unlike the usual probabilistic outputs of an LLM. Generating code to help in your data journalism feels like a specific use of AI that stands distinct from many others and I have far fewer concerns about it. To be clear, that’s only when you consider learning data journalism as the goal, not if your goal is learning to code; that’s an entirely separate discussion.
The Generative AI in the Newsroom blog from Diakopolous and Gilert at Northwestern University is a great live reflection of what the industry is trying out Regular posts from global newsrooms catalog the results of experimentation with AI; I read every single one. Their recent Agentic AI Investigative Challenge dove even deeper; the winners all showcased prototype investigative AI tools that centered provenance and reliability. These journalists are all trying things out and reflecting out loud about how it went.
Getting your hands dirty with tools, like authors for that blog do, is central to understanding their capabilities. I ask learners I work with to try things out because that is the best way to understand the opportunities and limitations of AI in data journalism.
Center Journalistic Principles
ProPublica’s AI principles page offer that they “treat AI-generated material as unverified source material.” Accountability , verification, transparency, integrity, privacy & security–their list is a pretty good set of core principles to guide your journalism. Sadly we have a growing number of AI gaffes in the daily operation of normal journalism practice that we can point to as the opposite, from the US state dept generating an incorrect map of Africa, to the New York Times correcting that a statement was an AI summary and not a quote of the Canadian opposition leader.
I take inspiration from well-written AI policy transparency like ProPublica’s, especially when it the ideas are so well-aligned with journalistic principles.
Readers see this. A study from Trusting News found that readers across 10 global news organizations overwhelmingly think AI disclosures are “very important” or “extremely important.” However, other research found that while readers value these disclosures, they still left the impression that AI connected journalism was less trustworthy. I have my students all turn in statements based on guidance from Trusting News’ “AI Trust Kit”. Writing these considered AI disclosures builds a strong reflective practice and reminds them that these core principles must guide all our work, no matter if it uses AI or not.
But how we develop these AI guidelines is also important. Most organizations start by looking at what AI is being used for and then consider what rules they need for those situations. This kind of approach, argues researcher Daniella DiPaola, is reactive and offers poor long term impact. Instead, she posits that we should start with what we care about most, and then develop AI policies that reinforce those core values. I find her values based approach to developing AI policies extremely helpful, and though it was developed for school systems I think it could be immensely valuable to newsrooms (and underlies how ProPublica presents their’s).
This kind of consideration of what principles guide engaging with AI in newsrooms came up as well in a recent Dutch thesis by van Bree. They describe how journalists engage AI tools in three ways: a colleague, an assistant, or an outsider. The colleague role cedes some authority and agency to AI while looking over it’s shoulder or helping it work. The assistant puts the human in charge of a collaboration with clearly defined tasks the AI would undertake. The outsider AI was seen primarily as a threat, devoid of verifiability and mostly needing attention as a subject of reporting. These three roles echoed my own experiences, where I move between considering what role the AI tools are playing. As noted, my CAReful approach means that I most often take the position that AI can be my assistant, following my directions and leaving the primary journalistic or design work of data journalism to me.
I appreciate van Bree’s naming of roles, and choose to highlight using AI as an assistant in my own learning and teaching. This is my own simple illustration.
Of course, this talk for SciCAR is in Germany, so I had to again laud the amazing Nazi party membership database project from Die Zeit as a well-considered inspiration in the investigative domain (which I admittedly know less well). They used AI up and down the chain of data processing for this deep investigative project. Throughout they offer strong examples of centering transparency, such as how they include an option for readers to offer a correction if some aspect of the processing of the hand-written cards led to an error or omission. Staff have shared with me that the story was a huge hit with readers, drove a major bump in subscribers and in turn created deeper investments within their newsrooms for AI investigative capabilities.
No matter what you learn or try with AI, centering journalistic principles is key. In the hype cycle and push to AI-ify everything we can’t lose sight of the core values that make data journalism what it is.
Stay Critical
My last guiding principle responds directly to the marketing and hype cycle of AI tools: we have to keep asking critical questions. In keeping with the rest of this piece, here I have more questions to offer as insights rather than answers. Sorry!
Is Big Tech a reliable partner? Remember the failed promises of a “pivot to video”? Offers of money, attention, and more from Big Tech are best viewed with skepticism. AI tools don’t cite news sources that they were trained on. AI company leaders don’t seem to view the misinformation and news-slop problem as critical. Our business models are not aligned. The recent history of journalism and Big Tech should make you constantly question if they are a reliable partner or provider of services.
Can we create alternatives? The AI moat might seem to large to cross, but if we can’t trust the California-centric monopoly perhaps we need to try. Model development is becoming more well understood, leading to public LLMs from countries from Switzerland, India, Poland, Spain and beyond. Even Reuters is getting in on the LLM-model building game, releasing their own “Thomson” model earlier this year. There are opportunities for custom AI work on specific tasks related to data journalism as well, as evidenced by the work Zetland did to build Good Tape, a revenue-generating transcription service now used by millions. Newsrooms and data journalism teams have long been an innovator in technologies (d3.js, Django, Svelte, etc.), A recent workshop I joined, led by my friend Harini Suresh at Brown, on “Collective Governance of AI in Newsrooms” left me hopeful that this energy is being directed at creating industry-led alternatives.
What of value is lost? Even when automation helps with tasks such as transcription, we have to consider the cost. Older journalists around me agree that the nuance of a speaker’s voice and repeated listening to a quote over and over build a subtle and deep relationship with a source when transcribing, one that is often critical to telling the right story. There are clear advantages to automated transcription, yet we also must introduce the loss of that relationship as a tradeoff to learners. In class I simply ask students to read a sentence, and then listen to the same thing read aloud from the source; the difference is obvious. This question is important to keep asking ourselves. For instance, what of value is lost when a robot is writing the d3.js code for your next interactive data visual?
Is AI good for learning? There is still much to see, but early evidence suggests that using AI can help in general assignments like homework assignments, but hurts in longer term learning like on exams. As I’ve already noted that I see a real gap in learning to code with AI; students who have a certain level of knowledge already can use and learn with AI coding tools quite quickly, but those who haven’t met that threshold learn significantly slower. Frank Elavsky writes about the misperception of how AI might help us wonderfully, nothing while we want AI to fill in skills needed to implement our wonderful ideas, in fact it is often used to produce high quality technical projects for poorly refined ideas. Conceptualizing and building are different and intertwined processes, requiring reflective use of tools to avoid the “slop-zone” he points to.
These four questions are the type of things I keep asking myself, and others I learn with, in regards to how AI might be used in data journalism. Asking ourselves hard questions is how we stay critical in this time.
Reflecting on Knowing
If you were expecting me to tell you to make charts with AI, or to not make chart with AI… sorry! These reflections are about the learning of data journalism in an AI era. If you know my background this won’t be surprising; I studied learning and computation for my M.S. at the Lifelong Kindergarten group at the MIT Media Lab. Yes, I went to kindergarten twice! The approach to designing creative learning experiences I learned there continues to inform all my teaching and research. More deeply, the idea of thinking about thinking continues to lead me to these kinds of moments of reflection.
It turns out that ability I was introduced to years ago is critical right now, because AI tools are changing how we think, learn and know things. The relevant academic term for this is “epistemology”: theories of how we know things. It’s a little meta but important to consider. Consider the simplistic example of ways you could know a soup is too salty: measuring the amount from bowl to compare to the recipe, tasting it yourself, hearing a trusted friend tell you, etc. I’m convinced that using AI tools to learn something changes how to you know it, though I think we don’t quite understand precisely how yet.
In older work I cited Heron’s concept of “extended epistemology” to highlight and explain that we have different ways of knowing things. I argue data journalist has a lot to learn from this.
This idea of how we know things matters in data journalism. In a talk for AoIR I argued that the off-screen arts-based approaches I work on in community data projects are a form of “epistemological resistance”, fighting the dominant formalized ways of knowing. This borrows from Turkle and Papert’s “epistemological pluralism” in software learning and theories of fighting colonialism from sociologist de Sousa Santos. AI changes how you think and learn in important ways, and you should be aware of that so you can resist it as needed. Reflection is an important part of learning, and one I’m trying ot practice right now with this piece.
Learn CAR-efully
So what does this mean for your learning and teaching? How do we apply my themes of being specific, trying things out, centering journalistic principles, and staying critical while AI technologies continue to be adapted into data journalism settings? Reflecting on their work in the Google / LSE JournalismAI initiative over the last few years, Schjøtt and Schaetz label what they saw as “domestication” of AI in newsrooms. They lay out four parallel paths the technology followed in newsrooms: demystifying AI to understand what it is, cultivating skills to use it, managing the limits, and appropriating it into practice. van Bree used the same term: “domestication.” If AI is a wild animal to tame in our data journalism, what methods are we to use?
I find those four paths of domestication useful to pause and consider for myself as a reflective learner. Sometimes you need to step away from the hustle and bustle of a tool and project to reconsider it from the ten-thousand foot level. In the end their report saw the processes led to both (1) increased confidence to try and engage with AI and (2) a growing sense of the inevitability of AI in the newsroom. There’s that tension again. That mixed outcome mirrors my own experiences on this space, and I’m sure will continue to complicate my under-construction approaches to teaching and learning AI in data journalism. Across those principles my key message at SciCAR is to learn CAR-efully. Hopefully the approaches and the examples I’ve shared inform and connect to your own experiences.
My closing side: learn CARefully.


