Graph Agents: Where it all started?
Sonam Bala
Author

Very rarely do we come across voice agents that do not break when the conversations venture into specific details with multiple turns involved. Be it a collections call where the agent gets the greetings right but forgets to verify identity despite having a cooperative callee on the other side, in a bid to confirm the due payment. This is observed even if there's a detailed prompt in the background that covers every single edge case possible.
Causation? The prompt was perhaps a kilometre long, and the agent kept scrambling for the right response and skipped parts despite the tokens getting maxed out. Getting lost while scrolling the source material right there, very humane of the agents, no?

The same collections call two ways: an instruction buried somewhere within one long prompt, and a graph where verification is a step the call has to pass through.
What about the agents that didn't break? What we saw in the past were mostly linear use cases where turns were few and the end goal was either a single-point agenda or just one task. The agents only broke when the variables in the call became too many to handle, and the agents began getting confused the moment the call was off-script. The fix? Perhaps reading the line step by step and doing the guesswork to get the RCA right. Weren't the agents built to make our jobs simpler?
Well, you have Graph Agents for that. They break the call into smaller chunks or steps that you can see and control before deploying.

One prompt carrying every step, against the same call split into five nodes.
An old idea with new applications never gets old
Breaking long, complex tasks into smaller, doable actionables isn't really a new concept. Veterans did it all the time, focussing on what was achievable for each step. Along came the idea of a graph, which is older than most of the software programs that it now facilitates.
The traces start with George Mealy and Edward Moore, who published the two state-machine models in '55 and '56. Named after them, the system resides in one state and transitions into another whenever a deterministic rule instructs it. That's also what happens on Bolna's voice agents, which have to handle more than one variable in a single call, impersonating humans for a better outcome.
During 1966-72, a few researchers at Stanford Research Institute came up with Shakey, a first-of-its-kind mobile robot that had displayed the capability of reasoning. Yes, this was more than half a century ago. It followed the A* search algorithm to move around, and STRIPS to plan its actionables. Every action was defined with a specific agenda and what changed, thus establishing that planning a route and its actions could all be compiled in one resident system.
What started from there soon followed its course to everyday software. The first industry that it disrupted was game design, where designers specifically used this to build behaviour trees and gamify the next move of any character in the game. Came 2014, Airbnb deployed graphs of dependent steps in Apache Airflow.
Soon, the ReAct paper, published in 2022, exhibited that language models could also reason, act, look at the results and follow the next steps accordingly. That's not AGI, but something if we look at how far we have come to establish how agents could be formed with determined values. Then in 2024, LangGraph and similar frameworks just standardised stateful graphs. The idea floated around for 70 long years by now, and found applications in different domains, only for an LLM to be placed inside one so that you could build an agent that truly listened to what your customers are saying, reasoned responsibly and responded accordingly, not just what you have fed the agent word by word.

A convergence of ideas over seventy years that lays the foundation for Graph Agents.
The problem voice has always had
When Voice AI became the default (or started getting widely accepted) as a mode of reaching out to people at scale, the tension was palpable. The industry had seen IVR menus, once standardised with VoiceXML in the early 2000s, which were divinely easy to audit, but rigidity got to them. Fixed responses on a few keys, while anything that wasn't covered would get only apologies.
With the onset of LLM agents, that became obsolete, with a high bar of reasoning for these conversations. The good agents had to have natural speech and code-switching, but having the instructions coded in a single, long prompt kind of lost the control that was once a given with IVRs. And in certain industries, the control wasn't optional but a mandatory compliance before deployment, e.g. lending, insurance or banking as a whole.
And now that we combine the same with the momentum of the deployments, voila, you have pressure. A study published in 2009 demonstrated that speaker behaviour was uniform across ten languages, i.e. nobody wanted overlapping turns or long silences, remarking latency. The usual gap sat around 200 ms, and every decision was awaited on an LLM to come in that would further add to the delay.

A classic IVR routes on keypresses. A graph agent lets the caller speak and routes on intent.
The solution: every single node in the flow gets one job
This is what Graph Agents solves for, precisely. It takes a problem statement with a lot of variables and splits the call into nodes, as many as required. Every node gets one dedicated job description and does precisely what it is instructed to do with defined exit rules. Every step, be it greetings, qualification, collection nudge or confirmations as well as closing, everything becomes a node without exceptions, with transitions as the thread joining them.
In the event of failure, the RCA points you to exactly what broke. You can change it without touching the rest. You can equate it with the structure of an IVR, and the caller continues talking as expected with the outcome you once desired, without having to depend on a human.
Sources
- Nilsson (ed.), Shakey the Robot, SRI Technical Note 323 (1984)
- IEEE Milestone: Shakey, the World's First Mobile Intelligent Robot
- Stivers et al., Universals and cultural variation in turn-taking in conversation, PNAS (2009)
- Levinson, Turn-taking in human communication, Trends in Cognitive Sciences (2016)
- Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models (2022)
- Apache Airflow project history
- W3C, Voice Extensible Markup Language (VoiceXML) 2.0 (2004)