Testing the Artistry Limits of AI Through The Kingdom series

July 22, 2026

Television series come in all forms and sizes. The Kingdom (1994), directed by Lars von Trier, is an absurdist, supernatural medical drama series that follows a ghost-hunting, medical malpractice, and plague plot in a hospital. Peter Leonard, an academic librarian and founder of the Yale Digital Humanities Lab, looks beyond the show’s plot and analyzes its frames and visual aesthetics. By using AI technologies, Leonard tests the limits of AI to see if it can move past recognizing basic objects and comprehend deeper, more abstract concepts in human artistry and cultural tension.

A presenter stands next to a lectern, with a large screen displaying a slide on the wall and another person speaking into a microphone.

Photo: Professor Tim Tangherlini introduces Peter Leonard and his talk to the audience.

Multimodal and Computational Analysis

The first test came with an old Convolutional Neural Network (CNN) from around the 2010s. Leonard wanted to see if a computer could group frames by how they look and, by using the semifinal layer of a CNN, 10,000 screen captures were gathered by visual similarity without facial recognition. Even though the CNN is an older technology, this proved that computers are great at finding visual patterns. Leonard noted that even a really old non-multimodal model is helpful in figuring out if, for example,, we can find all the scenes that are in the operating room. The next would test if the computer could find the right pictures based solely on a typed word. The model utilized was CLIP (Contrastive Language-Image Pre-Training). By typing words like “elevator” or “coffee”, all the screen captures with the object mentioned successfully showed up. Leonard also explicitly ran CLIP and typed the Danish word for coffee, getting the same images. “To prove to you that this is actually a multilingual model..." Leonard stated, noting that the model had learned to recognize concepts and objects across languages.

 Peter Leonard talking behind the podium, behind him is a screen presenting a slide showing the results of a multimodal model scene comparison.

Photo: Leonard showcases and compares similar scenes in visuality from the TV show The Kingdom.

Are computers able to understand why a TV show scene is interesting or strange? Understanding human artistry and culture sounds like a hard test for a computer. Leonard challenges this by showing the models mPLUG-Owl, Qwen2.5 VL, and the newer Qwen3 VL, classic films. The mPLUG-Owl model seemed to detect the weirdness in a scene showing a boy in a school uniform sitting in the middle of a desert, pointing out that it was a weird place to find the outfit. This is a useful result by the standards of 2026, demonstrating that even back then, with models from 2024, visual comprehension seemed to be a great ability of these models. However, Leonard noticed that by giving the Qwen3 VL model a prompt like “this is a scary hospital show”, the AI will utilize its built-in bias and try to please the user by calling every single scene “scary” even if it isn’t. “It was the sycophantic bias”, Leonard recalled.

Additional tests included checking if computers could automatically cut a long video into separate, clean story scenes. The models used this time were PySceneDetect and a deep-learning tool called TransNet V2. Unfortunately, this failed, as the direction and scene cuts came in chaotic artistic instructions, confusing the AI model and making Leonard cut all 52 scenes by hand. Here, Leonard highlights that the architecture simply struggles with narrative and looks more like a frame sampler. Another test included tracking characters switching back and forth between speaking Danish and Swedish. For this, OpenAI’s Whisper and NVIDIA Parakeet were utilized. These models struggled. Once a character started talking in one language, it would assume the whole clip was in that language. Switching languages mid-sentence got the AI confused, causing it to analyze the rest of the speech using the old language’s alphabet. An interesting problem, as Leonard describes it, arises as AI appears to be fundamentally not built for mid-phrasal code-switching.

A Look Into The Future

Peter Leonard concluded that while the modern multimodal model era shows great potential, studying complex cultural texts requires manually selected corpora and fine-tuned models capable of detecting cultural contexts, irony, and subtle linguistic variation, rather than relying entirely on generic and automated solutions. Inspired by an idea from director Lars von Trier, Leonard took a look into the future of digital humanities at the end of his presentation by remarking that "We have to take the good with the evil when it comes to multiple models."

If you are interested in this presentation, the full recording of Peter Leonard’s talk is available on the I School website. To stay informed and participate in discussions of critical cultural importance, we invite you to join our Cultural Analytics mailing list. You can subscribe by visiting our Google Groups page or by emailing bids-cultural-analytics+subscribe@lists.berkeley.edu.