Taking an AI to Art School: Building a RAG System for a VFX Curriculum

October 01, 2026

Building an AI assistant for a visual effects education platform presents a unique challenge: how do you enable an LLM to access thousands of hours of specialized video content, live instructor sessions, and large reference manuals? Slapping a ChatGPT widget on the app and saying “here’s your AI” wouldn’t work for anything beyond the basics, since the knowledge required is so specialized; fine-tuning a model would be slow and costly, since the curriculum material changes at the speed of technology. We settled on a Retrieval-Augmented Generation (RAG) system as the most flexible, powerful, and cost-effective option.

The goal of the AI assistant (“Synapse”) is to be the first line of response for students with questions or issues with the course material. By having Synapse handle simpler issues for the students, we can reserve the weekly instructor-led live sessions for more complex issues that require an instructor’s insight. Also, since the assistant is available 24/7/365, students in different time zones no longer need to wait for an instructor to become available to answer questions – they can get answers instantly.

There are three main content streams we wanted to include: pre-recorded course videos, live sessions, and reference manuals / policy documents. The course videos are the primary curriculum material; each week, the student receives a set of new lessons for that week, spread across 5 to 10 courses. Each lesson is made up of a set of video chapters, the audio of which is transcribed to text. The 90-minute live sessions occur weekly, and provide students with feedback and assistance from their instructor. Students are split into groups of about 8 for these sessions, and the same instructor runs the sessions for the group throughout the term.

Importing each of these streams presented a unique challenge: for the course video stream, the material changes regularly, so all material needs to be connected to the terms that the student is actually in. Live sessions are fairly unstructured and contain a lot of noise. Reference manuals and documents all have different formats and structures, and can be huge: a ‘small’ manual comes in at around 700 pages, while a larger one can reach 5000+ pages.

Course Material

For the course material, I fed each chapter in each lesson through an enrichment pipeline first; this pipeline uses a low-cost LLM (Haiku-level) to create a descriptive title for each chapter (replacing any “Lesson 01 / Chapter 01”-type titles), pulls out keywords, determines which software is being used in the chapter, and categorizes the chapter by type (“intro”, “main content”, “recap”, “Q&A”, and so on). After this, it chunks and embeds the transcript text, and saves it in the database with metadata attached.

Live Sessions

For live sessions, just dumping the transcripts into a database would yield a low-quality dataset: they contain a ton of social chat, “looks great; keep doing what you’re doing” feedback, and references to visuals that you can’t see. Instead of embedding the raw transcripts, I used an LLM to extract pairs of questions and answers from the meeting transcript, each of which must address a specific, well-described issue that the student is having. In the same step, I had the LLM remove any personal information from the text – not just first and last names (which were replaced by “the student” and “the instructor”), but any company names, email addresses, Discord usernames, etc. Because of the importance of sanitizing this information, I used a higher-end model (Sonnet-level) for this pipeline. From a starting set of around 500 meetings, I was able to extract roughly 5000 Q&A pairs, so about ten usable pairs per 90-minute meeting.

Manuals and documents

Software manuals and policy documents both have the same issue: each source has its own structure, format, and organization. To overcome this, I developed a custom parser for each document we wanted to add; this parser restructured the document into a standardized Markdown format, which was then run through the chunking and embedding pipeline. Splitting the documents into sections, then chunking these sections into roughly 400-token chunks, one by one, ensured that there was no cross-section overlap – each chunk contains information only relevant to its section, and the query would never return part of one section glued to another. I also curated the documents to ensure that we’re only importing information that would be actually useful (we don’t need to import change logs, for example, when only the current version of a policy is relevant).

Metadata

All of the pipelines for tag every embedding with as much metadata as possible: learning path, term, course, term week, and so on. This is crucial for getting high-quality results later on, as it reduces the surface area we need to search for each query – just hard-filtering by learning path can reduce the number of chapters we need to include by a factor of 4. We can also soft-filter by prioritizing metadata that meets certain criteria: for example, when querying live sessions, the live sessions for the same week in the previous term get a 3x bump in priority, on the assumption that students will likely have similar issues with a project that other students did in the previous term. In other words, if the student is asking about the Week 8 assignment, we should prioritize material from Week 8 in the previous term.

Query flow

With all the data and metadata in the RAG system, we now need to make it available to the end user. The app has an Agent Chat component which can call the LLM service; for each query, the LLM is provided with a set of search tools, from which it can select the ones appropriate for that query. After running the tools and retrieving the results, the LLM constructs a natural-language response from them, adds links to any specific videos in the results, and returns the response to the user.

Contextual Indicators

The RAG system also uses contextual clues to narrow and prioritize student searches based on metadata: if the student is inside a certain lesson, the chapters for that lesson get a priority boost; same idea if the student is inside a course or an assignment. If no relevant results are found, we can widen the search, since a student may be asking about a different course or lesson; but chances are good that the question is about the current lesson/course/assignment. This allows us to focus the initial search without explicitly asking the user “what are you looking for”.

Keeping it up-to-date

With a dynamic and ever-changing curriculum, the data from the initial pipeline run is only useful for so long. To keep everything up to date, the system uses a set of cron jobs to check for newly-uploaded videos, get their transcripts, and run them through the enrichment / embedding pipeline. If an admin copies a lesson or chapter thereof to a new course, the existing embeddings are copied to a new record, with updated metadata to reflect the new course and term. When new Zoom transcripts for live sessions become available, a cron extracts Q&A pairs from them, and adds their embeddings to the RAG database.

Lessons learned

  • Everything idempotent – any stage of the pipeline should be able to run multiple times, without worrying about duplicating data. This allows for graceful recovery from pipeline faults (like hitting the embedding API’s rate limit), and prevents costly re-running of data that we already have sent through the pipeline.
  • Metadata is key – composing metadata tags for the actual data should make up a significant part of the RAG system’s design and development. Being able to filter and prioritize results when it comes to actually querying the data gives you a massive relevancy bump.
  • Custom chunking – when source documents vary so much in structure, it’s worth developing a unique parsing strategy for each source. Resolving all documents to a common structure, then segmenting and chunking that intermediate document, gives you a much better chance of producing useful and accurate chunks than just feeding the original document straight into a chunking algorithm.

In a RAG system, the rockstar LLM gets all the attention, but most of the essential work happens upstream: cleaning, structuring, and tagging the data before it ever gets embedded

All content copyright James Podles, 2026

Not a robot — just an English major