Computational approaches in discourse analysis I

Marion Walton

Do you need a GenAI “assistant” for content analysis?

Outline

  1. What is discourse?
  2. Discourse analysis & Corpus linguistics
  3. Automating content analysis in Media studies
  4. Machine Learning & LLMs in media research

Quiz

  • Multiple Choice Questions (readings & lectures)

Objectives

  • Disciplinary roots of automated text analysis: Computational Linguistics, Natural Language Processing, Corpus Linguistics
  • Key concepts: corpus, frequency, collocation, keyness, concordance.
  • Useful complement to Critical Discourse Analysis

Questions

  • What disciplinary assumptions about language underlie various tools?
  • What is gained and lost when we automate text analysis?
  • How is text data prepared for analysis?

Tools for text analysis

Getting started

  1. Wordtree by Jason Davies
  2. Multilingual concordancer by Voyant Tools

Examples created with 4CAT,Orange, R language and the quanteda package which offers many useful text analysis functions.

Readings

  • Boumans, J. W., & Trilling, D. (2016). Taking stock of the toolkit: An overview of relevant automated content analysis approaches and techniques for digital journalism scholars. Digital Journalism, 4(1), 8–23. doi:10.1080/21670811.2015.1096598

  • Acosta, Fátima Ávila. 2025, Nov 10. AI in Social Sciences: How Large Language Models are Reshaping Text Analysis. The Algorithmic Review. https://www.globalcenter.ai/the- algorithmic-review/ai-in-social-sciences-how-large- language-models-are-reshaping-text-analysi

Additional

  • Baker,P.et al.(2008). A useful methodological synergy? Combining critical discourse analysis and corpus linguistics to examine discourses of refugees an d asylum seekers in the UK press. Discourse & society. 19.3.273–306.

  • Baker, P. (2006). Using Corpora in Discourse Analysis. Continuum: London. (Chapter on Concordance)

Example data

We will start with a relatively simple example (speeches) and then move to Clicks comments and metadata.

Some Quiz questions will focus on these datasets.

Mandela speeches

Mandela speeches

We will investigate the text from two historical speeches by Nelson Mandela. You can download them from the links below:

Clicks Videos

  1. Video descriptions and comments on YouTube videos posted about the Clicks/Tresseme controversy (2020). (see Amathuba for link to files)

What is discourse?

  • Discourse can be a problematic term - used in many ways.
  • Focus has been on language and written text but also other signifiers in context.

Linguistics

Discourse as the conventions which govern language "above the sentence" e.g. generic conventions in writing, and norms which shape interaction between speaker & addressee in spoken discourse

e.g. academic discourse, legal discourse, media discourse

Poststructuralism

Discourse as power and normativity, or “practices which systematically form the objects of which they speak” - e.g. medical discourse structures knowledge and social practices around health, constitutes entities e.g. “mental illness”, “homosexuality” and subject positions (doctor and patient) (see Foucault, 1972:49)

Media

“The hidden power of media discourse and the capacity of … power-holders to exercise this power depend on systematic tendencies in news reporting and other media activities.”(Fairclough, 1989:54)

Language as discourse

Every object or concept is surrounded by different ways of constructing it, which reflect different ways of representing the world. (Baker, 2006)

Language vs discourse

Language is just one part of discourse

Apart from using written language as a semiotic system, discourse can also be multimodal (visual, audio, haptic signifiers)

Discourse in Media research

Classic qualitative approaches to media research such as Critical Discourse Analysis (CDA) are traditionally focused on relatively small amounts of text.

Such approaches have continued relevance and new uses if they can scale up to engage with larger collections of text e.g. evaluating new computational tools and auditing biases and ideological assumptions in AI training data.

Can be expanded and enhanced by using computational approaches such as automated text analysis.

Interrogating AI

We need to know where new tools come from, or their provenance and what they assume about language and the world (e.g. what counts as “context”, English dominance, built-in gender biases etc).

Disciplinary roots

Computational linguistics, Natural Language Processing and Corpus Linguistics are related areas which provide different approaches, concepts and tools for analysing textual data.

Natural language processing

Natural Language Processing (NLP) is a subfield of Computer Science NLP develops algorithms and models for computers to “understand”, interpret, and generate human language in a contextually relevant way.

From Eliza the “therapy bot” to ChatGPT

Joseph Weizenbaum’s “DOCTOR”/Eliza , created in 1966 mimicked discourse patterns of an Rogerian therapist.

Meet Eliza, 1966

eliza

The Eliza effect

Anthropomorphic designs exploit the fact that users assume that computer systems which fluently output human language can also understand, think and engage in meaning-making as we do. This is known as the “ELIZA effect”.

NLP and media studies

Automatic text analysis methods used in media studies include:

  • Supervised text classification
  • Topic modelling
  • Word embeddings

Very useful for describing large collections of textual data such as social media posts, web pages, or interview transcripts.

Text as data

1964

Mandela’s speech from the dock, Rivonia Trial, 20 April 1964

1994

Mandela’s inaugural speech as President, 10 May 1994

Pronouns and nationalism

"I/me" highlights individual stance.

"We/us" groups people together, suggests community, masks power.

"They/them" Disidentifies, “others”

What is a Corpus?

A corpus is a set of documents which stores large quantities of real-life text. The plural form of the word is corpora.

You can find a set of South African language corpora on the SADILAR corpus portal website.

Individual documents or posts which make up the corpus can be labelled and stored separately from one another in the corpus format.

Frequency

Frequency is a key concept underpinning the analysis of text and corpora.

Nonetheless, as a purely quantitative measure it needs to be used with a sensitivity to

  1. The word-distribution patterns in human languages.

  2. The importance of context for meaning

Frequency & ideology

  • People have choices “No terms are neutral. Choice of words expresses an ideological position”. (Stubbs, 1996:107; Baker, 47)
  • “If people speak or write in an unexpected way, or make one linguistic choice over another, more obvious one, then that reveals something about their intentions, whether conscious or not.” (Baker:48)
  • Only certain choices are available at any one time (e.g. historically)

Problems with frequency

  • Can be reductive and generalising
  • Can oversimplify
  • Focus on differences between word distributions can obscure more interesting interpretations.

Exercise

Download the text of the Rivonia speech, copy it and paste it into the Wordtree tool and explore the patterns

Exercise

Use Wordtree to explore the Mandela speech from the dock (Rivonia Trial, 1964). How is Mandela’s pronoun use different to that in his inauguration speech as President in 1994?

Beyond frequency

Meanings are contextual

Need to go beyond simply counting frequencies.

Context plays an important role in meaning.

For discourse analysis, frequent clusters of words are more revealing than just looking at individual words in isolation

Concepts

  • Concordances
  • Collocation (next lecture)
  • Keyness (not studying keyness this year)

Concordance

Concordance

  • A list of all the occurrences of a particular search term in a corpus, presented within the context that they occur in, usually a few words to the left and right of the search term.

  • Co-text allows analyst to infer (some) context

Context allows us to address qualitative research questions

KWIC

  • Concordance is also known as “key word in context” or KWIC analysis

Starting points - Concordance

  • Open the local Voyant tools in your browser.
  • In September 2020, Clicks faced a public outcry about a TRESemmé advertisement posted on its website. How did YouTube videos describe the Tresemme advertisement? Use Voyant Tools to compare the concordances for the words controversial and racist in the video descriptions
  • Download and unzip the text files nlp_datasets_wordtree_voyant if you want to try using Voyant with the other corpora

Conclusion

Limits of “co-text”

Tools such as collocation, concordance and keyness allow us to investigate textual “co-text” and infer some contextual features.

Situated meanings are elusive but broader textual patterns can be distinguished.

Multimodality is a central aspect of context. Multimodal meanings are challenging to access with current tools.

Lecture 2 - Machine learning and GenAI

Machine learning

Most contemporary uses of the term “AI” actually refer to Machine Learning. Machine Learning is a subfield of AI which identifies patterns in data in order to use them for prediction. The word “learning” is used to refer to the way these systems can produce outputs which have not been explicitly programmed. Learning algorithms can be supervised or unsupervised. Supervised algorithms learn by example. Unsupervised learning algorithms classify data into groups of similar items.

Social media corpus

In this lecture we’ll be investigating a dataset of 200 Youtube videos focused on the controversy about a Tresseme advert posted on the Clicks website in September 2020.

The corpus includes:

  • Text descriptions of the videos (n=200)
  • Comments (n=3417) on a random sample of the videos (n=60)
  • AI Transcriptions of the audio tracks (n=58)
  • Video thumbnails (n=131) and keyframes (n=3267) from videos

ML & Clicks Video Descriptions

Framing

“Framing essentially involves selection and salience. To frame is to select some aspects of a perceived reality and make them more salient in a communicating text, in such a way as to promote a particular problem, definition, causal interpretation, moral evaluation, and/or treatment recommendation for the item described.” (Entman, 1993:52)

Machine learning and frame analysis

  • Co-occurrence of words interpreted as a frame using statistical techniques like cluster analysis.
  • Cluster analysis - organising items into groups, or clusters, on the basis of how closely associated they are.
  • Co-occurrences of words graphically visualized as networks of words

Proximity

We describe proximity of words using terms from linguistics

  • Collocation - the above-chance frequent co-occurrence of two words within a pre-determined span, usually five words on either side of the word under investigation. Finds “advert was not racist”, “advert was racist” and “racist advert”

  • Co-occurence - Above-chance frequency of ordered occurrence of two adjacent terms in a text corpus, indicator of semantic proximity (closeness of meaning) or an idiomatic expression e.g. “bad hair”

ML & Clicks Controversy

Collocations

We will explore collocation in the descriptions of the videos sample (n=200)

Top Collocations: Video Descriptions (n=200)
collocation count
julius malema 94
eff protest 21
across country 13
dry damaged 14
controversial advert 11
health beauty 8
stores across 12
black women 10
EFF members 14
racist advert 9
comment share 6
seed oil 6
face pack 18
following controversial 5
beauty retailer 5
thank much 5
tune afrika 8
subscribing liking 5
today eff 7
fine flat 6

Comments - Collocations

Top Collocations: Clicks Comments (n=35)
collocation count
2 black people 135
1 black women 106
3 white people 102
7 natural hair 48
9 black woman 45
4 fine flat 41
145 black hair 39
13 flat hair 35
5 dry damaged 28
15 can use 28
6 face cream 25
8 go back 25
27 black person 24
21 straight hair 22
49 can get 21
22 aneeza gold 20
76 white women 20
64 damaged hair 17
10 political party 16
11 thank ninja 16

Overview - Automated Content Analysis

Topic Modeling - News

News researcher can automatically categorize a huge collection of news articles into different topics. For instance:

  • Topic 1: Politics: “election”, “government”, “policy”, “vote”.
  • Topic 2: Sports: “game”, “team”, “score”, “tournament”.
  • Topic 3: Technology: “innovation”, “software”, “internet”, “startup”.
  • Topic 4: Health: “medical”, “treatment”, “healthcare”, “disease”.

Now the researcher can organize their content more effectively and find relevant articles based on their research question.

Topic Modeling - Comments

Topic modeling doesn’t work as well with small samples. For example Clicks comments data doesn’t divide neatly into separate topics.

      Topic 1  Topic 2     Topic 3     Topic 4   Topic 5  Topic 6  Topic 7 
 [1,] "clicks" "🤣"        "black"     "racist"  "just"   "can"    "hair"  
 [2,] "eff"    "go"        "people"    "us"      "like"   "one"    "white" 
 [3,] "must"   "know"      "think"     "country" "get"    "use"    "ad"    
 [4,] "right"  "say"       "women"     "time"    "need"   "also"   "even"  
 [5,] "take"   "u"         "😂"        "people"  "want"   "yes"    "look"  
 [6,] "malema" "see"       "racism"    "eff"     "love"   "good"   "make"  
 [7,] "thing"  "said"      "person"    "never"   "way"    "skin"   "fine"  
 [8,] "matter" "really"    "beautiful" "racism"  "going"  "face"   "saying"
 [9,] "stupid" "something" "white"     "now"     "well"   "please" "flat"  
[10,] "stores" "back"      "many"      "race"    "better" "thank"  "racist"

Topic Modeling

In NLP, topic modeling applies unsupervised learning on a corpus to produce a summary sets of terms representing the collection’s overall primary set of topics.

Topic Modeling

Used to identify topics present in a corpus.

LDA algorithm identifies co-occurrence patterns of words and latent structure of the text

LDA assumes: - each doc is a mixture of topics - each topic has characteristic word distribution

Why Use Topic Modeling

  • Exploratory work on large dataset
  • Summarise key themes
  • Reduces complexity and size of dataset (dimensionality reduction)
  • Faster Information Retrieval - find by themes not keyword matches

Goals of ML

The goal of machine learning is to:

  • make accurate predictions.
  • use large datasets
  • use complex models which recognise nonlinear relationships between several variables.

Boumans & Trilling, 2016

Boumans, J. W., & Trilling, D. (2016). Taking stock of the toolkit: an overview of relevant automated content analysis approaches and techniques for digital journalism scholars. Digital Journalism, 4(1), 8-23. https://doi.org/10.1080/21670811.2015.1096598

Approaches

Simple automation

Dictionary approaches (e.g. sentiment analysis)

  • specifies explicit rules
  • best for coding manifest data
  • sentiments are contextual - does not always travel well

“Sentiment” in Clicks transcripts

Supervised machine learning

Model "learns" from (encodes) decisions by human coders.

  • Makes more efficient use of human work.
  • Classifiers can be re-used and allow for faster response (e.g., Hopkins and King 2010; Jurka et al. 2013).
  • Inter-cultural generalisability?

Unsupervised machine learning

  • Unsupervised machine learning helps to describe discourses, frames or topics in an open way - doesn’t impose prior assumptions
  • Similar to qualitative methods
  • Difficult to audit
  • Can replicate bias

Strengths

  • Flexibility & scale
  • Classification tasks
  • Reality-based tasks with right and wrong answers

Weaknesses

  • Culturally constructed categories
  • Potential linguistic/culture/race/gender biases
  • Unintended uses
  • Variability & difficulties with auditing

Generative AI

Generative AI (GenAI) refers to computational systems used in popular chatbots (e.g. ChatGPT, Gemini, Claude, Copilot) which include generative models such as LLMs or diffusion models. Generative AI systems respond to prompts from users as input, which can include text and various forms of digital media. The systems produce synthetic media output (such as digital text, simulated dialogue, images, videos, audio, and software code) using generative models, which are statistical encodings of their training data.

GenAI in Research

How and to what extent should we use AI for academic work?

GenAI tools can improve efficiency for researchers, but they also carry risks, particularly the danger of overestimating the veracity and impartiality of AI-generated information.

Human oversight needed to avoid inaccurate or fraudulent results

(Acosta, 2025)

Strengths

LLMs excel at tasks like detecting emotions, identifying opinions, recognizing hate speech, and detecting topics

“zero shot” analyses (no training data required)

(Acosta, 2025)

Developing NLP scripts for reproducible analysis (MW)

Useful

The LLM answers short targeted questions about text content for text classification and categorization

e.g. Study of GBV news coverage - did each article explicitly mention the perpetrator or not?

(Acosta, 2025)

Straightforward

Using this thumbnail image from the YouTube video, transcribe any text in the image. Concatenate it with the text from [body] [videoTitle] and [channelTitle]

drawing

Challenging

Use the image and text together to classify the type of YouTube channel, assigning one best-fitting label from the following list:

PARTY_AFFILIATED: “Channels who are created by or endorsed by a South African political party. ” e.g. Economic Freedom Fighters, Cyril Ramaphosa, Democratic Alliance spokesperson, Member of Parliament for Action SA”

NEWS_CREATOR “Accounts who distribute and create news related content primarily through social and video networks. These actors are independent from wider news institutions. e.g. Infotainment, podcasters, comedy, gaming and lifestyle, commentary.”

TRADITIONAL_NEWS_ORGANISATION “Accounts which are affiliated with a legitimate mainstream media organization.” e.g Broadcasters, networks, news publishers, featuring professional journalists, reporters and clips from the accounts of these media organizations”

MUSIC_CREATOR: “Channel of a minor music production company or of a digital creator who creates and shares music and content related to music. For example, Pasto Espirito, Kay Cee, Sfatalito & Taa Nats, First Klass Terorrists”

OTHER: “Channel which does not obviously belong to any of the above categories”

drawing

Result

Channel types and publication dates in YouTube search results for “Clicks advert South Africa” 2020-2023

Validation

Any LLM classification needs careful validation and review

Example

drawing

“Former DA MP Renaldo Gouws joins FF Plus” - Timeslive 25 September 2025.

[Google] Gemini 3.1 Flash Lite: The channel belongs to Renaldo Gouws, a political commentator and content creator who provides independent commentary on South African news and political events. The video content focuses on reporting and critiquing the actions of a political party, which aligns with the definition of an independent news and commentary creator.

Example

drawing

Kat White is a journalist who sometimes posts stories on YouTube

[Google] Gemini 3.1 Flash Lite: PARTY_AFFILIATED: The content features individuals wearing EFF-branded clothing and explicitly promoting the party’s political actions, such as the protest against Clicks. This indicates the channel is directly aligned with and promoting the agenda of the Economic Freedom Fighters political party.

Example

drawing

Beauty influencer recommends a skin lightening product

[Google] Gemini 3.5 Flash: *BEAUTY_CREATOR Reason:** The transcribed text (“Lighten your skin safely with this 2 lightening lotion”) combined with the video description (“Use any of this 2 Lightening body lotion to lighten your skin safely for best result and Morexskinglow”) focuses on skincare, beauty tips, and product recommendations. The presenter in the thumbnail is a beauty influencer/creator sharing advice on skin lightening and achieving a “skin glow,” which aligns perfectly with the definition of a BEAUTY_CREATOR.

Concerns about GenAI

  • LLM classification needs human supervision and data validation
  • Limited applicability for low-resource languages (languages which are not well-represented online or in dictionaries, word lists, NLP tools etc.)
  • Pay careful attention to prompt quality
  • Challenges with reproducibility
  • Technical skills and costs
  • Representational and environmental harm [MW]
  • Privacy and data protection [MW]

(Acosta, 2025)

Conclusion

With careful validation, Generative AI can be helpful for low-level, focused tasks in content classification, unless data includes low resource languages.

Nonetheless, “AI-generated output alone is not the end goal for social scientists, but a starting point to explain broader social phenomena.” (Acosta, 2025)

Generative AI lacks understanding and moral agency. As a result it has limited capacity to engage with and interpret social realities, particularly in the Global South.

Importance of Critical Discourse Analysis

  • Essential perspectives if we are to address racism, sexism and Northern/Anglocentric biases in tools
  • Important skillset for auditing and adapting tools and datasets

Extra

Clicks data

You can download both datasets from the links below:

Further reading

  • Jurka, Timothy P., Loren Collingwood, Amber Boydstun, Emiliano Grossman, and Wouter van Atteveldt. 2013. “RTextTools : A Supervised Learning Package for Text Classification.”The R Journal 5 (1): 6–12.

  • Hopkins, Daniel J. & Gary King. 2010. A Method of Automated Nonparametric Content: Analysis for Social Science. American Journal of Political Science, Vol. 54.1. 2010. 229–247.

  • Vlieger, Esther, and Loet Leydesdorff. 2012. “Content Analysis and the Measurement of Meaning: The Visualization of Frames in Collections of Messages.” In Research Methodologies, Innovations and Philosophies in Systems Engineering and Information Systems, edited by Manuel Mora, Ovsei Gelman, Anette Steenkamp and Manesh S. Raisinghani, 322–340. Hershey, PA: Information Science Reference.

Models make a difference

YouTube #abahambe - xenophobic/afrophobic build up to March and March “deadline” June 30 2026

Classified using different models

School kids admitted in hospital

Qwen3 VL Flash

The thumbnail uses sensationalist language (“School Kids Admitted in Hospital”) combined with what appears to be coded political messaging. The phrase “Omayifrendi bayaqeda ngathi” (Omayifrendi has poisoned the food) is a known slogan used by the Economic Freedom Fighters (EFF) to accuse political opponents of poisoning food supplies — often as part of their anti-establishment rhetoric. The inclusion of “Ngcebo Biyela” (a real EFF politician) and the visual of men in traditional attire (possibly EFF supporters) reinforces this. The overall tone, imagery, and use of politically charged slogans strongly indicate this is a party-affiliated channel, likely promoting or reporting on content aligned with the EFF’s narrative. Label: PARTY_AFFILIATED

Gemini 3.5 Flash

NEWS_CREATOR Reason: The channel (Ngcebo Biyela) functions as an independent news creator and commentator. It distributes and discusses current affairs, social issues, and political commentary (specifically focusing on local controversies such as food safety in spaza shops and populist movements) directly to social media audiences, independent of traditional mainstream media institutions.