Our conversations at BIDS around AI often begin not only with what new tools can do, but also how these tools might reshape the questions researchers ask. On March 17th, as part of the Cultural Analytics workshop week, Peter Broadwell, Head of AI Modeling and Inference in Research Data Services at Stanford University Libraries, explored this question in his talk “Gesamtkunstvektoren: Perceptual Embeddings for the Performing Arts”. The title combines Gesamtkunstwerk, the idea of a “total work of art,” most closely associated with Richard Wagner, that gains its value from the interplay of its multiple artistic components, with Vektoren, the German word for vectors. Broadwell asked whether new multimodal embedding models can help researchers search, compare, and support the interpretation of performance recordings whose meaning emerges from the interaction of text, image, video, and audio. His talk showed how multimodal AI tools can make layered performing arts easier to search and analyze, while also emphasizing the significance of humanistic expertise to interpret the computational results.

Photo: Peter Broadwell describes multimodal models to the audience.
Multimodal models enable search across media formats
Broadwell began his talk by describing early models like CLIP (Contrastive Language-Image Pre-training), which allowed text-to-image search without needing captions or metadata about the images. He regarded CLIP as an early proof that aligning text and image understanding was powerful and practical. Next, he described more powerful multimodal embedding models, including Meta’s Perception Encoder Audiovisual (PE-AV) family, which can generate aligned semantic embeddings, meaning they represent video, audio, and language in a shared computational space so researchers can compare them with one another. As he put it, these models are “best used for indexing and for searching.” These tools can help researchers search for images with text, compare sound and video, or locate specific moments across a large audiovisual archive. This is especially useful for libraries, archives, and cultural heritage collections, where materials often exist across many formats. Moreover, for researchers working with performing arts, these tools can make complex recordings easier to navigate and explore.
Multimodal similarity does not equal artistic meaning
Pushing the models beyond what they are already good at, Broadwell took the audience on a journey to test whether these models could do something more interpretive. Instead of only using the models to search for objects or scenes, he tried to determine whether or not they could reveal “semantic resonances” among the layered elements of performance, including what is sung, seen, and heard. He chose to apply this to his earlier work on pose, action, and manually annotated intermediality in recorded theater and Noh performance, a traditional form of Japanese theater performed in masks and costumes that combines music, movement, chant, and gesture.
Broadwell used an example from the opera Nixon in China to illustrate how models could help people search large archives across text, visuals, video, and other formats. He compared audio, video, and text embeddings across different clips to test whether the model could correctly capture correlations of different parts of the performance. The models found some interesting correlations, but missed moments that would feel obvious to a human viewer. For example, Broadwell expected a strong correlation when the singer was holding the Little Red Book while singing “book,” but the model did not capture that moment well. This was an important reminder that computational similarity is not the same as cultural meaning. A model may surface patterns, but those patterns are incomplete and still need human interpretation.

Photo: Broadwell compares a scene from Nixon in China with a visualization of multimodal similarity scores.
In answer to a question posed by an attendee, Broadwell said contrastive models are useful because they do not give definitive answers; instead, they show correlations. If researchers want to ask richer questions about unity, intermediality, or performance, they need better ways of comparing model outputs and connecting them to humanistic interpretation.
Understanding patterns in cultural materials requires expertise
Broadwell’s talk demonstrated that cultural analytics is about more than just applying algorithms to cultural objects. It is about testing methods, understanding their limits, and bringing technical work into the conversation. Researchers can use AI as a research tool and apply the historical, artistic, and domain-specific knowledge needed to interpret the results.
If interested in this talk, please find the full recording of Peter Broadwell’s talk on the I School website. To stay in touch and join conversations of critical cultural importance, please join our Cultural Analytics mailing list by visiting the google groups page or emailing bids-cultural-analytics+subscribe@lists.berkeley.edu.