In ProgressSocial ImpactAI/LLMRAGNext.js 14

YojanaKhoj — AI Government Benefits Finder

Matching Indian citizens to 101 central and state welfare schemes in 3 minutes — on a rule engine where eligibility is true, false, or explicitly “cannot evaluate”, and the LLM explains the verdict rather than deciding it.

101
Schemes, 14 States
3-valued
Rule Logic
10
UI languages
help_outline

The Problem

Every year, ₹1.7 lakh crore in government welfare schemes go unclaimed in India. Not because people don't qualify — but because they don't know they qualify. 60%+ of eligible citizens never claim Ayushman Bharat. MGNREGA has ₹8,000+ crore in unpaid wages due to application failures. There are 400+ central and state schemes, spread across 30+ separate portals with different logins, different jargon, and different application processes. The people who need benefits most — BPL families, small farmers, rural women, senior citizens — are the least equipped to navigate this system. A 60-year-old farmer in Rajasthan shouldn't need to know what "PM-KISAN" stands for to receive the ₹6,000/year he legally qualifies for. The government's own solution — myscheme.gov.in — is a static directory. It lists all schemes, not *your* schemes. There is no personalisation, no guidance, and no plain-language explanation.
lightbulb

The Solution

YojanaKhoj (योजना खोज — "scheme search") is a conversational AI platform that asks 10–12 plain-language questions about a user's life situation, matches their profile against a database of 101 central and state schemes, and returns a personalised benefits report. The core user flow: → 3-minute conversational quiz (one question at a time, branching logic) → "Matching your profile..." — hybrid rule + AI eligibility engine runs in 3–5 seconds → Results page: schemes ranked by confidence (Definitely Qualifies / Likely Qualifies) → Each scheme: what you get, how to apply step-by-step, your document checklist, where to go → Download PDF report or share on WhatsApp The one-sentence pitch: "Tell us about yourself, we tell you every rupee the government owes you."
warning

Why It Was Hard

The naive approach — ask GPT-4o "does this user qualify for PM-KISAN?" — fails at scale for three reasons: 1. Cost. Every scheme × N users × GPT-4o calls = API bill that kills the product before it helps anyone. At ₹0.02–0.05 per LLM call, you need to be surgical about when you invoke the model. 2. Inconsistency. Government eligibility rules have hard boundaries (land holding ≤ 2 hectares) and soft edge cases (what if the wife owns the land, not the husband?). Pure LLM matching is inconsistent on the hard cases and the only reasonable approach for the soft ones. 3. Data quality. Government scheme documents are PDFs with inconsistent formatting, contradictory clauses, and outdated information. Building a reliable database meant AI-assisted extraction plus manual verification. The distribution problem is equally hard: my target users — rural, low-literacy, low-data — don't discover things via Google. The primary channel is WhatsApp. A website-first strategy, however well-built, misses 80% of the target population.
architecture

Architecture

The central design decision is that eligibility rules evaluate to three values, not two. `evaluateRule(rule, profile) → true | false | null` → true — the user passes this rule. → false — the user definitely fails it. → null — the rule cannot be evaluated, because the profile does not contain the field it asks about. Two-valued logic has no way to express "I don't know", so it has to lie in one direction. Treat unknown as false and you silently hide schemes from people who probably qualify — the exact failure the product exists to fix, and one nobody would ever report, because you cannot notice a scheme you were never shown. Treat unknown as true and you promise benefits that get refused at the counter, which is worse: a rural user has spent a day and a bus fare to find out. Making `null` a real value forces every consumer to decide what to do with it explicitly. During the hard filter, a scheme survives if all its hard rules return true or null — unknown gets the benefit of the doubt at the shortlisting stage, and the gap is then surfaced to the user as something to confirm rather than resolved silently on their behalf. The other decision that follows from it: the LLM never decides eligibility. It cannot. Hard rules are evaluated deterministically before any model is invoked. A scheme that fails a hard rule is gone — it is not in the list the LLM receives, so no amount of plausible reasoning can put it back. What the model does is rank the surviving candidates and explain them: two sentences in Hindi and English on why this scheme fits this person, what they are missing, and what to do next. Eligibility is arithmetic over rules; explanation is language. Only the second is the model's job. That split is also what makes the degraded path safe. When the LLM call fails, the system falls back to the top candidates by rule score with generic explanations — the verdicts are unchanged, only the prose is worse. Tech stack: Next.js 14 (App Router) + Node.js/Express on AWS Lambda + MongoDB Atlas + OpenAI GPT-4o + LangChain. Quiz engine: branching question flow stored in MongoDB. Each question has a followUpLogic map — if user says "farmer", next question is "land holding"; if "student", next is "education level". The backend tracks session state and serves the correct next question per answer. Derived boolean fields (has_bank_account, has_land_records) are computed from the documents the user reported, so rules can reference them without the quiz asking a separate yes/no question per document.
bug_report

What Failed First

First version sent every scheme through GPT-4o for every user profile. API latency was 25–40 seconds on the results page. Switched to rule-based pre-filtering: the full catalogue → ~20 candidates via deterministic rules → a top-10 shortlist sent to GPT-4o for ranking and explanation. Latency dropped to 3–5 seconds. The version before three-valued logic was the more instructive failure. Rules returned plain booleans, and a missing profile field evaluated to false. Test profiles that skipped the optional land-records question lost every farming scheme — including ones they clearly qualified for. Nothing errored; the results page was simply shorter. That is what pushed `null` into the return type: an unknown had to be representable, or it would keep being silently rounded to "no". Second failure: The quiz's branching logic was hardcoded in the frontend. When we needed to add a new question branch for a new state-specific scheme, it required a frontend deploy. Moved all branching logic to the database — each question document stores its followUpLogic map. New branches via admin panel, no code change. Third issue: PDF report generation was running synchronously. For users with 15+ matched schemes, the PDF took 8–12 seconds to build while the user stared at a spinner. Moved report generation to a background Lambda function — the user sees results immediately, PDF generates async, download button activates when ready. The biggest unresolved problem is distribution. The WhatsApp bot — the primary channel for rural users — is built: handler, templates, and webhook are in the repo behind a `WHATSAPP_LIVE` kill-switch, waiting on Meta WhatsApp Business API approval and BSP template sign-off. Until that clears, the live channel is the website, which serves literate, urban users. The engineering is done; the gate is external.
insights

Results (Current Status: In Progress)

MVP is functional end-to-end: quiz → AI matching → results → scheme detail → PDF download. Current coverage: • 101 schemes seeded — 16 central and 85 state schemes across 14 states • UI in 10 languages (en, hi, bn, te, mr, ta, kn, gu, pa, or); scheme content in Hindi and English — every scheme record carries paired _hi and _en fields, so explanations are generated in both rather than translated after the fact • Three-valued rule engine + LLM ranking running at 3–5s latency • PDF report generation • Admin panel for scheme management • WhatsApp bot built and gated behind a kill-switch, pending BSP approval What's left before launch: • Broaden state coverage beyond the current 14 states • Flip WhatsApp live once Meta/BSP template approval clears • Application tracker — let users mark which schemes they applied for • Alerts for new schemes matching a saved profile • NGO dashboard for bulk beneficiary management The product is technically ready to deploy. The distribution problem — reaching rural users who won't find a website — is the unsolved challenge that matters most, and it is now waiting on an approval rather than on code.
school

Key Learnings

1. Model the unknown explicitly. A rule engine over incomplete user data needs three values, not two — the moment "cannot evaluate" collapses into "false", you start hiding outcomes from exactly the people you built the thing for, and you get no signal that it is happening. 2. Decide with rules, explain with the model. Eligibility is arithmetic over stated criteria; a model that can be talked into a verdict has no business producing one. Give the LLM a shortlist that is already correct and let it do the part it is actually good at — saying why, in the user's language. 3. Distribution is harder than the product for civic tech. The people this helps most won't Google it. WhatsApp-first, website-second is the right architecture for rural India — not an afterthought. 4. Data quality is the real moat. Anyone can build the AI layer. Building and maintaining a verified, structured database of government schemes with up-to-date eligibility rules, documents, and portal links — that's the defensible asset. 5. Build the admin panel on day one. A scheme database that requires a developer to update is a scheme database that gets stale in 2 weeks. The admin panel for non-technical scheme managers was not an afterthought — it was a launch requirement. 6. The user you're designing for can't give you feedback. Low-literacy, rural users won't file a GitHub issue. Testing with proxy users (NGO field workers, CSC operators) gave more actionable feedback than any analytics dashboard.