back to portfolio
// case study · GenAI · RAG

How NyayaAI reached 0.93 Recall@5

Building a citation-grounded RAG assistant for the Constitution of India: structure-aware chunking, hybrid search with reranking, and evaluation gates that decide what ships.

Shubham Kumar Gupta · source on GitHub

Legal questions punish a RAG system that is only roughly right. If someone asks about the right to education, an answer that cites Article 21 instead of Article 21-A is wrong, however fluent it sounds. So I built NyayaAI around three ideas: measure retrieval, handle failure explicitly, and attribute every claim to an Article.

NyayaAI architecture: the PDF is chunked and stored in pgvector; questions go through the chat API, router, hybrid search with reranker and answer LLM
One Postgres, no agent loop, no LangChain retrievers.

01Chunk by the document's own structure

Fixed-size windows cut Articles in half and mix neighbouring provisions. Instead, NyayaAI makes one chunk per Article and splits only long Articles, at clause boundaries. All 506 of 506 Articles in the Contents list are found, with no duplicates or gaps, and no chunk exceeds the 1,024-token limit. A full ingest takes about two minutes, and re-ingesting the same PDF is a no-op.

The real PDF had its share of surprises, and each fix is pinned by a test:

02Hybrid search, then rerank

Each query runs a dense vector search (bge-m3 embeddings in pgvector) and a Postgres full-text search. The two lists are fused with Reciprocal Rank Fusion, and a cross-encoder (bge-reranker-v2-m3) reorders the candidates. Exact lookups like "Art. 21-A" are pinned, so they never depend on similarity scores.

The ablation shows which piece earns its keep:

VariantRecall@5MRR@10nDCG@5Cand. recall@15
Dense only0.8890.8460.8020.972
Full-text only0.6250.5630.5160.486
Hybrid (RRF)0.7780.7070.6640.875
Hybrid + rerank (current)0.9310.9310.8770.875
Dense + rerank0.9580.9620.9110.972

The reranker adds the most, taking Recall@5 from 0.778 to 0.931. The honest surprise is the last row: on this set, dense + rerank beats the current hybrid setup. Whether to keep the full-text leg is an open decision, recorded in an ADR rather than settled by gut feel.

03Evaluation gates, not vibes

A golden Q&A set with a dev/test split drives everything. make eval-retrieval reports Recall@k, MRR, nDCG, candidate recall and latency, and the gates block merges:

MetricValueGateStatus
Recall@50.931≥ 0.90pass
MRR@100.931≥ 0.75pass
Hit@1 (Article lookups)1.000= 1.00pass
Retrieval latency p9595 ms≤ 300 mspass
Candidate recall@150.875≥ 0.95open
Rerank latency p954,846 ms≤ 800 msopen

Two gates still fail, so the baseline isn't promoted yet. Writing that down matters more than a headline number. The misses are instructive too: "supreme commander of the armed forces" misses Article 53 because the text says "supreme command of the Defence Forces", and Article 300A gets crowded out by 31A–31D.

04One router call, one answer call

NyayaAI is a fixed LangGraph pipeline, not an agent. A small LLM rewrites the question, extracts Article references and picks a route (article_lookup, simple, conceptual, multi_part, or a canned reply for ambiguous, out-of-scope and chit-chat messages). A large LLM then answers only from the retrieved, numbered excerpts. Conceptual questions use a HyDE passage for the vector search, and multi-part questions run one search per sub-question.

05Guards that keep answers honest

Takeaway: the biggest gains came from the reranker and from respecting the document's structure. The most useful habit was letting the evaluation harness, not intuition, decide what stays.

View the code Get in touch