Home/LangGraph fan-out
Fan out the agents. Cap how many run.
A batch of questions becomes one LangGraph agent per question. They run in parallel, five at a time, so a long list cannot fire every model call in the same instant. The next batch waits on a throttle and a concurrency slot.
- 01
One batch arrives
research-fanout starts as a single graph run. ThrottlePolicy allows 20 of these starts per minute.
- 02
Agents fan out in slices
.map() launches one agent per question. The loop keeps five of those agents in flight, then takes the next five.
- 03
The next batch waits
ConcurrencyPolicy keeps five graph runs in flight. A sixth waits up to three minutes for a free slot.
Paste this pipeline
agent_node wraps a compiled LangGraph graph as a normal node, so .map() fans it out. The reason step below is a stand-in. Swap it for ChatOpenAI when you want real tokens. The fan-out and the limits stay put.
from graphingest import graph, deploy, ThrottlePolicy, ConcurrencyPolicy
from graphingest.langgraph import agent_node, AgentConfig
def build_researcher(config: AgentConfig):
from langgraph.graph import StateGraph, END
from typing import TypedDict
class State(TypedDict):
messages: list
def reason(state: State) -> dict:
question = state["messages"][-1]["content"]
# Replace this with ChatOpenAI(model=config.model).invoke(...)
answer = f"[{config.model}] notes on: {question}"
return {"messages": [{"role": "assistant", "content": answer}]}
builder = StateGraph(State)
builder.add_node("reason", reason)
builder.set_entry_point("reason")
builder.add_edge("reason", END)
return builder.compile()
researcher = agent_node(
name="researcher",
graph_builder=build_researcher,
config=AgentConfig(model="gpt-4o-mini", max_iterations=6, stream_steps=False),
)
@graph(
name="research-fanout",
timeout_seconds=3600,
throttle=ThrottlePolicy(limit=20, period_seconds=60),
concurrency=ConcurrencyPolicy(limit=5, wait_timeout_seconds=180),
)
def research_fanout(queries: list[str]):
width = 5 # agents in flight inside this run
results = []
for start in range(0, len(queries), width):
results.extend(researcher.map(queries[start : start + width]))
return results
if __name__ == "__main__":
deploy()width = 5is the cap inside one run. A 500-question list still finishes, in slices of five.ThrottlePolicy(limit=20, period_seconds=60)is the cap across runs: 20 starts of this graph per minute.ConcurrencyPolicy(limit=5, wait_timeout_seconds=180)keeps five runs going. Another run waits for a slot.
Run it
pip install "graphingest[langgraph]"- Run
sdk/python/examples/langgraph_fanout.py.deploy()registers the graph, then the three sample questions fan out. - Watch the run on the dashboard. Each agent is its own task, so a failed question retries without redoing the ones that finished.
Other jobs
Vercel cron timeouts
The route returns in under a second. The pipeline runs for hours.
Slack when a job fails
A message in the morning, with a link to the run.
GitHub Actions
GitHub starts the job. The mark updates when the work finishes.
TypeScript and Go
deploy() onto managed infrastructure. Idle cost is zero.