From YouTube Captions to Q&A: Build Your Own Video Knowledge Assistant
A video knowledge assistant that turns YouTube captions into a searchable question-and-answer workflow using retrieval and language models.

Representational image
While watching some business podcasts, I noticed an important gap: when a key concept like aggregate demand comes up, who is responsible for breaking it down in context? And what if there are more such questions, as is often the case? This led me to think—why not combine video transcripts with an LLM to query and explore these concepts in detail? That idea became the foundation of this project.
Wouldn’t it be better to just ask the video directly without watching it entirely?
“Explain aggregate demand in the context of this video?”
“Provide a gist of this video.”
This project shows how you can turn any YouTube video into a searchable knowledge base.
One Python script does it all: fetch captions → clean them with LLMs → index them → chat interactively.
🔧 Tech Stack
- Python 3.11+—use pyenv if you are not on 3.11. I had not yet tested the pipeline on other versions.
- youtube-transcript-api—fetch captions
- OpenAI API—caption cleanup and optional Q&A answers. I chose an external LLM because the pricing for
gpt-4o-miniwas attractive. A cost-aware system estimates how much you are using. - rank_bm25—keyword relevance (sparse search)
- sentence-transformers—semantic embeddings (dense search)
- Hybrid Retrieval—BM25 + Embeddings combined for better results
- dotenv + logging—config & cost tracking
- (Optional) Docker—you can containerize, but up to you.
📂 Code Structure
This project is just one script (yt_caption_pipeline.py) with clear modular sections:
- Environment Loader → reads .env or OS env vars. You have to create this manually. Structure details at the script. Feel free to tweak for variable responses.
- Transcript Fetcher → pulls captions from YouTube
- Cleaner → LLM removes noise, smooths text, tracks token costs
- Indexer → segments transcript, stores BM25 + embeddings
- Hybrid Searcher → combines keyword + semantic scores
- Interactive Shell → ask questions, view answers, exit with final cost
Everything runs in sequence, no external DB needed.
🌀 Flow Diagram

Script flow
🚀 Usage
- Clone the repo, instructions for sparse cloning below.
git clone --depth 1 --filter=blob:none --sparse https://github.com/pain459/projects/
cd projects
git sparse-checkout set tube_desc
cd tube_desc
- Create a virtual environment
- Install the dependencies. These are the versions I used:
numpy==2.3.2
openai==1.101.0
python-dotenv==1.1.1
sentence-transformers==5.1.0
tiktoken==0.11.0
youtube-transcript-api==1.2.2
- Run the pipeline
python yt_caption_pipeline.py --url "https://www.youtube.com/watch?v=abcd1234" --lang en --prefer any --llm-answer
- Example run:
(.venv) ravik@FED1:~/src_git/projects/tube_desc$ python yt_caption_pipeline.py --url "https://www.youtube.com/watch?v=eHJnEHyyN1Y" --lang en --prefer any --llm-answer
21:12:33 | INFO | .env loaded via dotenv from: /home/ravik/src_git/projects/tube_desc/.env
21:12:33 | INFO | Fetching transcript for video: eHJnEHyyN1Y (lang=en, prefer=any)
21:12:34 | INFO | Fetched transcript with 340 segments; ~14194 characters.
21:12:35 | INFO | Split into 3 chunk(s). per_chunk≈1500, overlap=150
21:12:35 | INFO | [1/3] Cleaning chunk (chars=5885)...
21:13:02 | INFO | [2/3] Cleaning chunk (chars=4350)...
21:13:25 | INFO | [3/3] Cleaning chunk (chars=4494)...
21:13:43 | INFO | Stitching chunk boundaries...
21:13:52 | INFO | Wrote cleaned transcript → cleaned_eHJnEHyyN1Y.txt
21:13:52 | INFO | Building index for eHJnEHyyN1Y from cleaned_eHJnEHyyN1Y.txt using all-MiniLM-L6-v2
21:13:52 | INFO | Segmented into 11 segments (size≈1200, overlap=200)
21:13:56 | INFO | Index written → store/eHJnEHyyN1Y
Model: gpt-4o-mini
Tokens (cleaning) in=8395 out=2770 cost≈$0.002921
Starting interactive search...
Search ready. Ask a question. Type 'quit' or 'exit' to finish.
Query> What is the gist of this video?
ANSWER:
The video discusses the entrepreneurial journey of Lynda Weinman, who created Lynda.com in 1995 to showcase graphic design and teach online. It highlights her counterconventional mindset, which contrasts with traditional business practices. The narrative also includes examples of other entrepreneurs, like Tris and Becs with Go Ape, who innovatively utilized existing resources, and Jonathan Thorne, who addressed specific surgical challenges. The video emphasizes the importance of focusing on real problems rather than merely product changes, showcasing how successful entrepreneurs identify niche markets and leverage available assets to build impactful businesses.
[Session tokens] in=9733 out=2884 cost≈$0.003190
Query> What are the key takeaways?
ANSWER:
Key takeaways include:
1. **Upfront Cash Strategy**: Tesla successfully sold 100 Roadsters for $100,000 each, generating $10 million in cash to fund production. This principle of securing upfront payments has been crucial for their growth.
2. **Innovative Business Models**: Go Ape leveraged existing resources by partnering with the Forestry Commission to create adventure sites, demonstrating resourcefulness in entrepreneurship.
3. **Counterconventional Mindsets**: Entrepreneurs often challenge traditional business practices. Examples include Arnold Correia's flexibility in business and Lynda Weinman's transition from teaching to online learning, which led to a billion-dollar sale.
4. **Emphasis on Cash Flow**: Entrepreneurs prioritize cash flow as essential for sustaining ventures, contrasting with large companies that may misallocate excess cash.
[Session tokens] in=11088 out=3044 cost≈$0.003490
Query> quit
Final session token usage → input: 11,088 | output: 3,044
Estimated total cost: $0.003490 USD
(.venv) ravik@FED1:~/src_git/projects/tube_desc$
🌟 Why This Matters
- Search smarter—hybrid retrieval gives you both precision (BM25) and meaning (embeddings).
- Stitching technique helps in reconstruction of LLM responses and returned in chunks.
- Stay cost-aware—token usage is tracked live.
- Lightweight—single script, no DB, no heavy dependencies.
- Extensible—add Docker, swap embedding models, or build a web UI later.
📌 Next Steps
- Add multi-video search (cross-video Q&A)
- Try local LLMs to cut API costs
- Wrap in a Streamlit/Gradio app
- Enhance summarization features while reporting token usage.
- Explore integrations with a knowledge base. It would be interesting to build a parallel brain that you can search and apply clustering algorithms to while maintaining diversity.
Originally published on Medium on August 26, 2025.