Ever searched for something and gotten back results that don't even contain the words you typed? That's not magic. There's a whole pipeline running underneath. Most people have heard the term RAG thrown around but have no real picture of what it actually does.
So I built something to show it. rag-playground is a recruiter search tool framed as a demo. You type a query like "find me a backend engineer with fintech experience", and instead of just showing you results, it shows you every step of how those results got found: the query turning into numbers, those numbers being compared against 50 candidate profiles, the closest ones getting pulled, and an LLM writing a recommendation from exactly that context, streamed back word by word.
Why keyword search isn't enough
Traditional search is exact match. Search for "backend engineer fintech" and you get documents containing those words. But what if a candidate's profile says payment infrastructure at scale instead? Keyword search misses them entirely.
Semantic search understands meaning rather than text. Both phrases end up near each other in vector space, so the right candidate still surfaces. That difference is immediately visible in the similarity scores the demo shows you.
RAG takes this further. Instead of asking an LLM to answer from its training data, you first fetch the most relevant content from your own database, then hand that to the model as context. The answer is grounded in real data rather than whatever the model happened to learn.
The pipeline
The demo runs every query through four stages, each one lighting up as it completes.
① Embed the query
The query text gets sent to OpenAI's text-embedding-3-small model, which converts it into a 768-number vector. Similar meanings end up close together in that space. The UI shows you the first 8 numbers as a preview so the concept doesn't stay abstract.
② Semantic search
That vector gets compared against precomputed embeddings for all 50 candidate profiles using cosine similarity in PostgreSQL via pgvector. The top 5 closest matches come back with a score between 0 and 1. This is the step where a profile about "payment systems" can rank above one that literally says "fintech" because the vectors are closer in meaning.
③ Context assembly
The top 5 profile bios get assembled into a structured prompt block. The UI shows you a preview of exactly what gets handed to the LLM, along with an estimated token count.
④ LLM generation
Gemini writes a recruiter recommendation using only the retrieved profiles as context. The response streams back token by token over WebSocket so you watch it build rather than waiting for a full response to drop.
Stack
| Layer | Choice |
|---|---|
| Backend | Go + Encore framework |
| Embeddings | OpenAI (text-embedding-3-small) |
| Completion | Google Gemini (gemini-2.5-flash) |
| Vector Search | PostgreSQL + pgvector with HNSW index |
| Realtime | WebSocket for pipeline events, SSE for token streaming |
| Frontend | Next.js App Router + Tailwind CSS |
| Deploy | Encore Cloud (backend) + Vercel (frontend) |
I kept vectors in PostgreSQL with pgvector rather than pulling in a dedicated vector database. For 50 profiles it's the right call and the HNSW index keeps similarity search fast. The architecture stays simple without giving anything up.
Built with Go, Encore, Next.js, PostgreSQL, pgvector, and Google Gemini.