Home/LangGraph fan-out

Fan out the agents. Cap how many run.

A batch of questions becomes one LangGraph agent per question. They run in parallel, five at a time, so a long list cannot fire every model call in the same instant. The next batch waits on a throttle and a concurrency slot.

  1. 01

    One batch arrives

    research-fanout starts as a single graph run. ThrottlePolicy allows 20 of these starts per minute.

  2. 02

    Agents fan out in slices

    .map() launches one agent per question. The loop keeps five of those agents in flight, then takes the next five.

  3. 03

    The next batch waits

    ConcurrencyPolicy keeps five graph runs in flight. A sixth waits up to three minutes for a free slot.

Paste this pipeline

agent_node wraps a compiled LangGraph graph as a normal node, so .map() fans it out. The reason step below is a stand-in. Swap it for ChatOpenAI when you want real tokens. The fan-out and the limits stay put.

from graphingest import graph, deploy, ThrottlePolicy, ConcurrencyPolicy
from graphingest.langgraph import agent_node, AgentConfig

def build_researcher(config: AgentConfig):
    from langgraph.graph import StateGraph, END
    from typing import TypedDict

    class State(TypedDict):
        messages: list

    def reason(state: State) -> dict:
        question = state["messages"][-1]["content"]
        # Replace this with ChatOpenAI(model=config.model).invoke(...)
        answer = f"[{config.model}] notes on: {question}"
        return {"messages": [{"role": "assistant", "content": answer}]}

    builder = StateGraph(State)
    builder.add_node("reason", reason)
    builder.set_entry_point("reason")
    builder.add_edge("reason", END)
    return builder.compile()

researcher = agent_node(
    name="researcher",
    graph_builder=build_researcher,
    config=AgentConfig(model="gpt-4o-mini", max_iterations=6, stream_steps=False),
)

@graph(
    name="research-fanout",
    timeout_seconds=3600,
    throttle=ThrottlePolicy(limit=20, period_seconds=60),
    concurrency=ConcurrencyPolicy(limit=5, wait_timeout_seconds=180),
)
def research_fanout(queries: list[str]):
    width = 5  # agents in flight inside this run
    results = []
    for start in range(0, len(queries), width):
        results.extend(researcher.map(queries[start : start + width]))
    return results

if __name__ == "__main__":
    deploy()
  • width = 5 is the cap inside one run. A 500-question list still finishes, in slices of five.
  • ThrottlePolicy(limit=20, period_seconds=60) is the cap across runs: 20 starts of this graph per minute.
  • ConcurrencyPolicy(limit=5, wait_timeout_seconds=180) keeps five runs going. Another run waits for a slot.

Run it

  1. pip install "graphingest[langgraph]"
  2. Run sdk/python/examples/langgraph_fanout.py. deploy() registers the graph, then the three sample questions fan out.
  3. Watch the run on the dashboard. Each agent is its own task, so a failed question retries without redoing the ones that finished.