__ ___ __ ___
/ \ \ / |__ |__) \ / | |__ | |
\__/ \/ |___ | \ \/ | |___ |/\|
Objective: Design a form that communicates the relationship between a specific text and a specific image, and which includes attribution among its objectives. You must use machine learning, either working with it or working against it.
This project unfolds over three phases. These phrases are intended break down the technical skills in smaller digestible chunks. Each phase should push your technical understanding forward. This stands in contrast to other assignments you may have had at Pratt, where phases are meant to give you milestones in your progression to the file product. You are expected to consider the high-level project objective as you work through the phases.
Each phase will serve conversation fodder for the following phase.
As we layer text into our mental model of AI technologies, we will focus on the relationship of text and image. There are several frameworks that theorize this relationship. Today the AI industry is focused on image-text coherence, which is an equivalence in signification between the two parts: the text describes the image, the image illustrates the text. Historically there have been other theories that leave room for each side’s independent function. There have also been frameworks that imagine a fusion of text and image fulfilling a function that neither can achieve on their own. Our sources will range from art history to semiotics to media theory. We will learn about and discuss how these frameworks compare and contrast.
Throughout this unit we will learn how these technologies actually work. We will use that technical knowledge to investigate how the technologies produce political effects. For example, we will have lectures that describe how machine learning has been criticized for re-encoding societal power structures and for the invisible labor required to train it.
This course in general intends to hone your sense of authorship in an AI-saturated landscape. These tools complicate traditional notions of authorship. While authorship has historically been associated with originality and effort, today it calls in concepts like collaboration, complicity, implication, and exploitation. We will relate this to this philosophical concepts like the death of author, and economic ones like invisible labor.
This assignment requires you to use machine learning in your process. That said, you do not need to know how to code. I will provide you sandboxed environments to experiment with two technologies.
The first is Word2Vec, which is a text embedding model that allows you explore how words in the training data relate to another. One of its main features is the ability to perform “algebra” with words. For example, this model’s numerical understanding of its training data “proves” that king - man + woman = queen . Students can access the sandbox here: https://word2vec.typo.school/
The other is CLIP which stands for Contrastive Language-Image Pre-training. CLIP is the foundation for our ability to prompt image generators with text. Its purpose is to provide a shared numeric understanding of an image and the text that describes it. In machine learning terms, CLIP has two encoders that produce vectors: one that encodes images, one that encodes text. CLIP’s job to make sure that for an image/text pairing, the image’s vector is numerically similar to the text’s vector. Students can access the sandbox here: https://image-text.typo.school . The password is ThinkLikeAComputer.
▄▄▖ ▗▖ ▗▖ ▗▄▖ ▗▄▄▖▗▄▄▄▖ ▗▄▄
▐▌ ▐▌▐▌ ▐▌▐▌ ▐▌▐▌ ▐▌ █
▐▛▀▘ ▐▛▀▜▌▐▛▀▜▌ ▝▀▚▖▐▛▀▀▘ █
▐▌ ▐▌ ▐▌▐▌ ▐▌▗▄▄▞▘▐▙▄▄▖ ▗▄█▄▖
Objective: Create two inter-related compositions (one typographic, one image-based) that explore Word2Vec arithmetic
This phase asks you to familiarize yourself with latent space, the idea that an embedding is a set of coordinates in a high-dimensional space, where each dimension is a particular aspect the machine can see. As numeric values, embeddings can be pushed in any one of those directions. Moving around this space can reveal the machine’s biases. But there’s a catch: these dimensions are not named, and they reflect only what the model was trained on.
Details
Background information
Relevant class materials
Recommended reading
Due date: October 6, 2026