At 11:35, I had a blank repo and one question:
Can voice give a student direction?
A few hours later, Disha could hold a real conversation in Hindi — and follow a student into Hinglish or Marathi mid-sentence — uncovering the constraints behind a career choice and speaking back with the warmth the moment deserved.
The hardest career question is rarely about careers.
A student says, “Mujhe engineering karni hai.” Underneath that sentence may be a fee ceiling, a parent who will not allow a hostel, a two-hour travel limit, or a dream borrowed from a cousin. A form captures the answer. A good counsellor discovers the reason.
Disha was built for students in tier-3 and tier-4 towns who speak the way India actually speaks: Hindi carrying most of the thought, with English and Marathi moving through it. Voice was not a feature. It was the only interface that made sense.
“The product had to ask the questions nobody had asked—and still feel like someone worth answering.”
Sarvam made the hardest part feel almost unfair.
In most voice products, the first hours disappear into plumbing: transcription, accents, language detection, and synthetic voices that flatten emotion. With Sarvam, the first loop worked early enough that I could spend the remaining time on the product’s judgment.
It treated code-mixing as speech, not noise.
Students could say “fees ka issue hai, hostel afford nahi kar payenge” without adapting themselves to the machine — and when one slipped into Marathi mid-sentence, that stayed intact too. Saaras kept the messy, useful meaning.
Explore Saaras v3 ↗It gave the interface a personality people could trust.
Bulbul did more than read text aloud. Indian names sounded right. English words inside Hindi did not break the rhythm. When Disha slowed down after a student mentioned family pressure, the voice carried the change.
Explore Bulbul v3 ↗Sarvam documents Saaras v3 for 22 Indian languages plus English, with code-mixed output; Bulbul v3 supports 11 languages and 30+ natural voices.
A conversation with a real state change.
The demo was not “ask AI anything.” Every layer had a job, a boundary and an observable outcome.
- 01ListenSaaras turns real, code-mixed speech into usable context.
- 02ReasonDisha asks one human question at a time and remembers the constraints.
- 03Ground327 pathways, 41 scholarships and two handbooks keep the advice honest.
- 04SpeakBulbul returns the answer in a voice that belongs in the conversation.
“A few hours” only works when you know what not to build.
- Idea lockedOne honest counselling conversation, in the student’s own language.
- Voice aliveHindi in. A natural Bulbul voice back. The riskiest layer was working.
- Product, not demoConstraints, grounded search, memory and a useful summary connected.
- Language followedA Marathi student stopped being answered in Hindi. Five turns, five in Marathi.
- Boundaries heldThe wellbeing stop moved out of the prompt and into the code that runs it.
- The room got loudBackground voices stopped counting as the student’s voice.
- She sounded like herselfHindi marks the speaker’s gender on the verb. A woman’s voice needed a woman’s grammar.
- ShippedImages built, deploy green, and the conversation held under test.
The last hour went to three things I did not expect.
The prompt lost. Disha is told, in plain words, that if a student says something frightening the career conversation is over. It flagged the moment correctly, gave the helpline, and then asked how far from home she could study. One turn later. Instructions are a request; the tools are the rule. The stop is now enforced where the model cannot talk its way past it — the career tools simply refuse for the rest of that session.
The room got loud. Browsers already suppress noise, but they suppress the wrong kind. A fan is easy. People talking nearby is not — it lives in the same frequencies as speech and arrives in the same syllable rhythm, so the agent kept hearing the room as an interruption and stopping mid-sentence. Cancelling background voices, not background noise, is what a real demo room needs.
“Two of the bugs I chased hardest were not in the product. They were in the ruler.”
A test run had scored the safety flow at 2 out of 5. The agent had done nothing wrong — the worker restarted sixty seconds into the call and the transcript recorded silence as failure. Another scenario reported the agent replying with fragments of its own interrupted sentences; the harness was reading the tail of the sentence it had just cut off and calling it the answer, shifting every reply one step. Both looked exactly like model failures. Neither was. Before believing a bad score, check that the thing measuring it was alive and pointed at the right moment.
Sarvam gave me leverage. Experience told me where to apply it.
Previous startup work taught me the invisible parts of voice: interruption handling, endpointing, recovering after a user changes their answer, and the difference between a fluent response and a useful one. I did not need to discover those failure modes during the build.
It also taught me to protect the core loop. No login system. No avatar theatre. No giant database. Just a phone, one tap, a trustworthy conversation and a summary a student could take home.
India-first AI should not feel like a translation.
Sarvam let me build from the student outward instead of from the model inward. The language could be mixed. The thought could be unfinished. The accent could be local. The product still listened.