Graph Agents: Where it all started?

Sonam Bala

Sonam Bala

Author

Graph Agents: Where it all started?
A telephone operator at her switchboard in 1922, routing each call down a cord you can see. Graph Agents bring that back to the voice call.

Very rarely do we come across voice agents that do not break when the conversations venture into specific details with multiple turns involved. Be it a collections call where the agent gets the greetings right but forgets to verify identity despite having a cooperative callee on the other side, in a bid to confirm the due payment. This is observed even if there's a detailed prompt in the background that covers every single edge case possible.

Causation? The prompt was perhaps a kilometre long, and the agent kept scrambling for the right response and skipped parts despite the tokens getting maxed out. Getting lost while scrolling the source material right there, very humane of the agents, no?

A long collections prompt where the line "Verify the caller before sharing any details" is skipped, next to a graph where greet, verify_identity, disclosure, due_amount and close are nodes the call must pass through in order

The same collections call two ways: an instruction buried somewhere within one long prompt, and a graph where verification is a step the call has to pass through.

What about the agents that didn't break? What we saw in the past were mostly linear use cases where turns were few and the end goal was either a single-point agenda or just one task. The agents only broke when the variables in the call became too many to handle, and the agents began getting confused the moment the call was off-script. The fix? Perhaps reading the line step by step and doing the guesswork to get the RCA right. Weren't the agents built to make our jobs simpler?

Well, you have Graph Agents for that. They break the call into smaller chunks or steps that you can see and control before deploying.

One long prompt, where every step competes for attention, splitting into five focused nodes: Greet, Verify, Collect, Confirm and Close

One prompt carrying every step, against the same call split into five nodes.

An old idea with new applications never gets old

Breaking long, complex tasks into smaller, doable actionables isn't really a new concept. Veterans did it all the time, focussing on what was achievable for each step. Along came the idea of a graph, which is older than most of the software programs that it now facilitates.

The traces start with George Mealy and Edward Moore, who published the two state-machine models in '55 and '56. Named after them, the system resides in one state and transitions into another whenever a deterministic rule instructs it. That's also what happens on Bolna's voice agents, which have to handle more than one variable in a single call, impersonating humans for a better outcome.

During 1966-72, a few researchers at Stanford Research Institute came up with Shakey, a first-of-its-kind mobile robot that had displayed the capability of reasoning. Yes, this was more than half a century ago. It followed the A* search algorithm to move around, and STRIPS to plan its actionables. Every action was defined with a specific agenda and what changed, thus establishing that planning a route and its actions could all be compiled in one resident system.

What started from there soon followed its course to everyday software. The first industry that it disrupted was game design, where designers specifically used this to build behaviour trees and gamify the next move of any character in the game. Came 2014, Airbnb deployed graphs of dependent steps in Apache Airflow.

Soon, the ReAct paper, published in 2022, exhibited that language models could also reason, act, look at the results and follow the next steps accordingly. That's not AGI, but something if we look at how far we have come to establish how agents could be formed with determined values. Then in 2024, LangGraph and similar frameworks just standardised stateful graphs. The idea floated around for 70 long years by now, and found applications in different domains, only for an LLM to be placed inside one so that you could build an agent that truly listened to what your customers are saying, reasoned responsibly and responded accordingly, not just what you have fed the agent word by word.

Timeline from Mealy and Moore machines in 1955 to 56, through Shakey, A* and STRIPS in 1966 to 72, behaviour trees in games in the 2000s, Apache Airflow in 2014, ReAct in 2022 and LangGraph in 2024, to Graph Agents on Bolna in 2026

A convergence of ideas over seventy years that lays the foundation for Graph Agents.

The problem voice has always had

When Voice AI became the default (or started getting widely accepted) as a mode of reaching out to people at scale, the tension was palpable. The industry had seen IVR menus, once standardised with VoiceXML in the early 2000s, which were divinely easy to audit, but rigidity got to them. Fixed responses on a few keys, while anything that wasn't covered would get only apologies.

With the onset of LLM agents, that became obsolete, with a high bar of reasoning for these conversations. The good agents had to have natural speech and code-switching, but having the instructions coded in a single, long prompt kind of lost the control that was once a given with IVRs. And in certain industries, the control wasn't optional but a mandatory compliance before deployment, e.g. lending, insurance or banking as a whole.

And now that we combine the same with the momentum of the deployments, voila, you have pressure. A study published in 2009 demonstrated that speaker behaviour was uniform across ten languages, i.e. nobody wanted overlapping turns or long silences, remarking latency. The usual gap sat around 200 ms, and every decision was awaited on an LLM to come in that would further add to the delay.

On the left, a classic IVR main menu routing "press 1" to Billing, "press 2" to Support and anything else to "Sorry, I didn't get that". On the right, a caller speaking freely and a graph agent routing by intent to a billing node or a support node

A classic IVR routes on keypresses. A graph agent lets the caller speak and routes on intent.

The solution: every single node in the flow gets one job

This is what Graph Agents solves for, precisely. It takes a problem statement with a lot of variables and splits the call into nodes, as many as required. Every node gets one dedicated job description and does precisely what it is instructed to do with defined exit rules. Every step, be it greetings, qualification, collection nudge or confirmations as well as closing, everything becomes a node without exceptions, with transitions as the thread joining them.

In the event of failure, the RCA points you to exactly what broke. You can change it without touching the rest. You can equate it with the structure of an IVR, and the caller continues talking as expected with the outcome you once desired, without having to depend on a human.

Sources