AI toolchain · proof of concept · public broadcasting

ARD Mediathek /
identifying moods and motivations and weighing them with AI

A proof of concept for ARD to find out if AI can generate two subjective metadata fields, mood and consumption motivation, reliably enough to support the editorial teams. A big part of the work was deciding how to measure this, because people don't agree on these categories either.

My role

Experience designer

Ownership

Google AI Studio toolchain · consistency metrics and analysis tool · insights · project lead in the evaluation phase

Team

One senior experience designer · one project manager · one developer on the Azure toolchain

Duration

3–4 months

ARD produces more than 2,000 hours of content every day, and each video has 23 editorial metadata fields. Editorial teams fill them in by hand, next to their actual work. Each team does it a bit differently, so the quality and completeness of the metadata varies a lot.

This has consequences on both sides. For viewers, content is harder to find and the recommendation algorithm works with weak data, which frustrates younger users in particular. For ARD, the manual work is an extra load on the editorial teams, and partner platforms with their own metadata requirements can only be served with considerable extra effort.

Our task was a proof of concept for two of the most difficult fields, Mood and Consumption Motivation, including an analysis of whether AI tools could take over part of this work.

2,000h

of ARD content produced per day

23

editorial metadata fields per video, and growing

2

subjective fields tested: Mood and Consumption Motivation

Mood and consumption motivation describe how a programme feels and why people watch it. There is no single correct answer, so before testing any AI we needed a structure to measure the results against. My colleague had started this work before I joined the project: 30 moods grouped into 5 clusters and 13 sub-clusters, and 27 consumption motivations based on ARD's own media research, grouped into 4 clusters with their sub-clusters. Clusters reduce the variance in the results, but they don't remove it.

For terms that are hard to tell apart, like quirky and absurd or provocative and controversial, there was a reference document we called the Mood-Bible. It has a definition, recognition criteria and examples for every mood, and decision trees for the borderline cases.

We also wrote down how we expected the AI to fail, and built a countermeasure into the prompt for each case. The model could add terms that are not in the list, so the rules only allow the given vocabulary. It could answer from its training data, for example from what it already knows about a famous series, so the prompt limits it to the data of the video itself. And it could invent an answer instead of showing uncertainty. For that reason every tag comes with a score, a confidence level and a justification anchored in the data, and at the end the model checks its own output against the rules.

Toolchain 1 · indirect

Video

image and audio

Azure Video Indexer

extracts summary, transcript, emotions, content moderation and labels

Claude Sonnet 4.5

interprets the extracted data, never sees the video

Tags

3–5 moods, 1–3 motivations, with scores

Toolchain 2 · direct

Video

image and audio

Gemini in Google AI Studio

reads image, audio and text in one pass, including light, colour and tone of voice

Tags

3–5 moods, 1–3 motivations, with scores

↑ The two toolchains. Both received the same prompt, Mood-Bible and rules

01

Two toolchains. We compared two different ways of getting from a video to tags. The first one is indirect. Azure Video Indexer extracts a summary, the transcript, emotions detected in the spoken text, content moderation signals with severity levels and labels for objects and scenes, and Claude Sonnet 4.5 interprets this data. Claude never sees the video, only what Azure extracted from it. The second one is direct: Gemini in Google AI Studio processes image, audio and text in one pass, so it can also use things like lighting, colour or the tone of a voice. My colleague worked on the Azure toolchain with support from a developer. I set up and ran the Google toolchain on my own, which was possible without technical resources because AI Studio doesn't require programming.

02

The prompt and the test set. The prompt guides the model through four steps. First it identifies 5–8 candidate moods. Then it scores each one with a decision tree: how central the mood is to the structure of the video sets a base range (60–69, 70–84 or 85–100), its presence in the text adds up to 15 points and the emotion data up to 10. After that it ranks the moods and keeps the 3–5 strongest with a score of at least 60. The last step checks for duplicates, scores that don't match their justification and a wrong order. One extra rule stops the model from picking several near-synonyms from the same sub-cluster, so the result describes the video more broadly. Consumption motivation works in a similar way, but starts with an analysis of the target group and returns 1–3 motivations marked as primary or secondary.

For the test we chose eight titles from different genres and formats: a crime drama, a feature film, several documentaries and reportages, a hybrid format and a children's show. Each title ran five times on each toolchain, 80 runs in total, and every run was exported as a CSV.

Snippet from the system prompt showing its four-step structure

↑ Snippet from the system prompt

03

Measuring consistency. When the runs were finished I took the lead on the project. Together with my senior colleague I worked out what we needed to know: across five runs, does the AI choose moods from the same clusters, from the same sub-clusters, the same terms, and how stable are the intensities? From these questions we defined three metrics. Each of them compares the five runs of a video in pairs and averages the result.

Presence consistency

How many tags two runs have in common, divided by the size of the larger set. Calculated at cluster, sub-cluster and term level.

Intensity stability

100 minus the average score difference for the tags both runs chose.

Position consistency

How far the shared tags move in the ranking from one run to the other.

Test criteria

Defined before testing. Consistency: in every run at least 3 of 5 tags from the same cluster, and at least 60% overlap in intensity. Quality: in 3 of 5 runs, at least 3 of 5 tags in the clusters chosen by the human panel.

The AI assistants we first used for these calculations gave different numbers every time, so I wrote a small Python script instead, with AI helping me with the parts of the code I didn't know. I had watched all eight videos and was also part of the human panel, which helped me notice when a result could not be right. It took a few iterations until the numbers matched what we had seen.

04

The human panel. As a control group, eight colleagues tagged the same videos, four people per video. They used the same vocabulary and the same score ranges, from 60–69 for a mood that is recognisable but not dominant to 90–100 for one that dominates the whole video. They worked without the decision trees and the Mood-Bible.

Both toolchains passed the consistency criteria for all eight videos. At cluster level, the five runs agreed on mood between 68 and 92% of the time. Intensity stability was above 91% for mood and above 89% for consumption motivation, and the ranking of the shared tags was stable as well. Lower values only appeared at term level, where the AI moved between neighbouring terms inside the same cluster.

The human panel agreed much less, between 36 and 54% at cluster level for mood. These numbers are not directly comparable, since one model repeating itself is not the same as four different people agreeing, and the panel worked without the decision trees. Still, it showed us that subjective metadata needs shared definitions and some training for the people doing it as well.

Table of mood presence consistency scores for both AI toolchains across the eight test videos

↑ Mood presence consistency, AI toolchains compared across the eight test videos

For quality we compared the tags of the AI with the clusters chosen by the panel. Mood passed for all eight videos on Google and for seven on Azure; the eighth, the hybrid format, passed at cluster level but not at term level. Consumption motivation was harder, and three videos stayed below the threshold on both toolchains. For Tatort and the children's show, the analysts had context knowledge the model didn't have. For a documentary about male body image, the AI saw orientation and help, while the analysts saw inspiration and curiosity.

Clusters are more stable than terms

The AI moved between neighbouring terms but stayed in the same cluster. Building search and recommendations on clusters and sub-clusters would make the results more reliable.

Context changes the answer

For many viewers Tatort is a Sunday evening ritual. Without that context the model tagged it as thrill-seeking, opinion-forming and moving. With it, the tags became something to talk about afterwards, a sense of comfort and opinion-forming.

Format matters

Google worked better with fragmented formats like the children's show, where the transcript alone loses the smaller segments. Azure and Claude worked better with the hybrid format, which also split the human panel: some read it as sensual, others as dramatic or frightening.

The rules mostly held

In 80 runs neither toolchain invented a term or a justification. The exception was Google using what it already knew about Tatort, although the prompt didn't allow it.

One conclusion changed how we framed the next phase. If people disagree with each other more than the model disagrees with itself, matching the human panel can't be the only definition of a correct result. The definitions themselves need to be validated first, ideally together with ARD's media research.

After I left, the team interviewed colleagues from ARD partner management, WDR's machine learning team and WDR's metadata competence unit, and sent questions by email to ARD's programme directorate and to ZDF Streaming.

"Wann seid ihr fertig? Ich könnte das gebrauchen."

"When will you be done? I could use this." · Interview participant, ARD

"Verschlagwortung sollte 100% autonom laufen, da kann man mit Fehlern umgehen."

"Keyword tagging should run fully autonomously, there we can live with mistakes." · Interview participant, ARD

80

AI runs analysed across 2 toolchains

68–92%

cluster consistency across five runs per video

MVP

next phase planned from the findings

The PoC was largely successful. Both toolchains were similarly strong, each with its own strengths, and neither invented terms or reasons. The findings went into a requirements list for the next phase, and several of the requirements come directly from the PoC: AI-generated metadata has to be marked as such so editors can review it, the output should include a confidence level, editors must be able to correct and complete the tags, and series and fragmented formats need their own rules. Age ratings were explicitly excluded from AI generation. By the time I left, the next phase was being planned as an MVP.

Tool selection→ AI-ready metadata→ Data foundation→ AI toolchain setup→ User testing→ Rollout
Next case study Solid Design System