10 ms·
Launch HN: Golpo (YC S25) – AI-generated explainer videos
Hey HN! We’re Shraman and Shreyas Kar, building Golpo (https://video.golpoai.com https://video.golpoai.com), an AI generator for whiteboard-style explainer videos, capable of creating videos from any document or prompt.
We’ve always made videos to communicate any concept and felt like it was the clearest way to communicate. But making good videos was time-consuming and tedious. It required planning, scripting, recording, editing, syncing voice with visuals. Even a 2-minute video could take hours.
AI video tools are impressive at generating cinematic scenes and flashy content, but struggle to explain a product demo, walk through a complex workflow, or teach a technical topic. People still spend hours making explainer videos manually because existing AI tools aren’t built for learning or clarity.
Our solution is Golpo. Our video generation engine generates time-aligned graphics with spoken narration that are good for onboarding, training, product walkthroughs, and education. It’s fast, scalable, and built from the ground up to help people understand complex ideas through simple storytelling.
Here’s a demo: https://www.youtube.com/watch?v=C_LGM0dEyDA#t=7 https://www.youtube.com/watch?v=C_LGM0dEyDA#t=7.
Golpo is built specifically for use cases involving explaining, learning, and onboarding. In our (obviously biased!) opinion, it feels authentic and engaging in a way no other AI video generator does.
Golpo can generate videos in over 190 languages. After it generates a video, you can fully customize its animations by just describing the changes you want to see in each motion graphic it generates in natural language.
It was challenging to get this to work! Initially, we used a code-generation approach with Manim, where we fine-tuned a language model to emit Python animation scripts directly from the input text. While promising for small examples, this quickly became brittle, and the generated code usually contained broken imports, unsupported transforms, and poor timing alignment between narration and visuals. Debugging and regenerating these scripts was often slower than creating them manually.
We also explored training a custom diffusion-based video model, but found it impractical for our needs. Diffusion could produce high-fidelity cinematic scenes, but generating coherent sequences beyond about 30 seconds was unreliable without complex stitching, making edits required regenerating large portions of the video, and visuals frequently drifted from the instructional intent, especially for abstract or technical topics. Also, we did not have the compute to scale this.
Existing state-of-the-art systems like Sora and Veo 3 face similar limitations: they are optimized for cinematic storytelling, not step-by-step educational content, and they lack both the deterministic control needed for time-aligned narration and the scalability for 5–10 minute explainers.
In the end, we took a different path of training a reinforcement learning agent to “draw” whiteboard strokes, step-by-step, optimized for clear, human-like explanations. This worked well because the action space was simple and the environment was not overly complex, allowing the agent to learn efficient, precise, and consistent drawing behaviors.
Here are some sample videos that Golpo generated:
https://www.youtube.com/watch?v=33xNoWHYZGA https://www.youtube.com/watch?v=33xNoWHYZGA (Whiteboard Gym - the tech behind Golpo itself)
https://www.youtube.com/watch?v=w_ZwKhptUqI https://www.youtube.com/watch?v=w_ZwKhptUqI (How do RNNs work?)
https://www.youtube.com/watch?v=RxFKo-2sWCM https://www.youtube.com/watch?v=RxFKo-2sWCM (function pointers in C)
https://golpo-podcast-inputs.s3.us-east-2.amazonaws.com/files/4c26c0cf-4938-4371-a74b-a78eb18acc86.mp4 https://golpo-podcast-inputs.s3.us-east-2.amazonaws.com/file... (basic intro to Gödel's theorem)
You can try Golpo here: https://video.golpoai.com https://video.golpoai.com, and we will set you up with 2 credits. We’d love your feedback, especially on what feels off, what you’d want to control, and how you might use it. Comments welcome!
- mandeepj 1y agoCongrats on the launch! If I may ask - how do you generate your audio?
- typs 1y agoIf that demo video is how it actually works, this is a pretty amazing technical feat. I’m definitely going to try this out. Edit: I've used. It's amazing. I'm going to be using this a lot.
- skar01 1y agoThank you!!
- Masih77 1y agoI call bs on training a RL agent to literally output strokes. The way each image renders is a dead give away that this is just using a text to image model, then convert it to svg, and finally animate the svg paths. They might even bypass the svg conversions with clever mask reveals. I was able to achieve the same thing in about 5 mins. https://giphy.com/gifs/rFVxSxZMlflZUX4TqI https://giphy.com/gifs/rFVxSxZMlflZUX4TqI
- mclau157 1y agoI have used AI in the past to learn a topic but by creating a GUI with input sliders and output that I can see how things change when I change parameters, this could work here where people can basically ask "what if x happens" and see the result which also makes them feel in control of the learning
- skar01 1y agoThank you!!
- skar01 1y agoHey also, if you want to suggest a video, we could try generating one and reply here with a link! Just tell us what you want the video to be about!!