Today’s project session was one long debugging arc: prove the agent has no memory, give it memory, then stop that memory from bankrupting us.
1. Proving the Problem
We compiled the existing graph as graph_v0 and rendered it as a Mermaid PNG with IPython.display to confirm the flow visually.
Then the live test:
- Turn 1: asked for a recommendation → the agent suggested the Radiancy Vitamin Cream.
- Turn 2: “which product did you just recommend?” → the agent denied recommending anything.
Printing the message history explained it: turn 2 didn’t contain turn 1’s messages. Without a checkpointer, every .invoke() is a brand-new session.
2. Checkpoints and Thread IDs
from langgraph.checkpoint.memory import MemorySaver
checkpointer = MemorySaver()
graph_v1 = state_graph.compile(checkpointer=checkpointer)
cfg = {"configurable": {"thread_id": "customer-1"}}
graph_v1.invoke({"messages": [...]}, cfg)
Two things, both required:
- The checkpointer saves state after each step. Note the keyword is
checkpointer, notcheckpoint. - The thread ID says whose conversation this is. It must sit inside
configurable.
Proof it worked: turn 2 went from 2 messages to 6. The history was being carried.
3. Inspecting State
snapshot = graph_v1.get_state(cfg)
snapshot.config # contains the checkpoint_id
snapshot.values # iterate to print each message
MemorySaver keeps all this in RAM. In production it belongs in a database.
4. Simulating Multiple Customers
We wrapped it in a chat function that reads input, exits on exit, quit or an empty line, and passes a thread ID through.
- Customer 1: a skin concern, quoted a serum at ₹699 for 30ml
- Customer 2: oily skin, recommended a cleanser at ₹349 for 100ml
Separate thread IDs kept the two conversations completely separate, and resuming an old thread ID picked the conversation back up.
5. Two Flaws We Found
Nodes returning the whole state. Inefficient, and it risks overwriting the message history. Return only what changed:
def assistant(state):
return {"messages": [reply]} # not the entire state
Nothing is persistent. MemorySaver dies with the process. Production needs SQLite or Postgres:
# pip install langgraph-checkpoint-postgres
from langgraph.checkpoint.postgres import PostgresSaver
with PostgresSaver.from_conn_string(DB_URI) as checkpointer:
checkpointer.setup() # creates tables on first run
graph = builder.compile(checkpointer=checkpointer)
6. The Cost Problem
The messages list grows with every turn, and you resend all of it each time. At 50+ turns the token bill stops being trivial.
Since it’s just a list, we can prune it. The plan: summarise once the conversation gets long. We accept slightly lower fidelity in exchange for bounded cost.
7. The Summarization Node
| Parameter | Value | Notes |
|---|---|---|
SUMMARY_TRIGGER | 12 messages | 20 was the first proposal; tune it |
KEEP_RECENT | 6 messages | could be 50–150 on a larger-context model |
| Summary length | under 120 words | may need to grow to ~500 in testing |
The prompt rules: keep the customer’s status, their concerns and the specific products discussed, and never invent anything.
State and continuity
- Added a
summarystring field to the advisor state. - Each new summary extends the previous one rather than replacing it, so nothing silently drops.
- The LLM call returns the summary; we extract just the content.
Cleaning up old messages
from langchain_core.messages import RemoveMessage
return {"messages": [RemoveMessage(id=m.id) for m in old_messages]}
A safe_cut_messages helper decides how many to drop while always preserving a human message, so the history never starts mid-thought. RemoveMessage only works when the state key uses the add_messages reducer.
Routing
route_after_assistant: tool calls present → the tools node; otherwise → the summarization route.route_exit: message count over the trigger → thesummarizenode.- We used matching node and route names to keep the conditional mapping simple.
8. Worth Comparing: The Built-In Option
LangGraph and LangMem already ship trim_messages, summarize_messages and a prebuilt SummarizationNode. Build yours first, then read theirs and see which decisions they made differently. That comparison is the real exercise.
✅ Your Assignment
Treat today’s code as a draft from a developer, and yourself as the tester.
- Clone it, run it, and debug the summarization flow.
- Any failure is a defect for you to fix, not a question to bring to class.
- Fix the LLM prompt so the model actually uses the
summaryfield. Right now it’s written but not consumed. - Test with real conversation data and find the word limit that works.
- Convert the nodes to partial state updates.
9. Still Open
- Structured extraction: customer name, mobile, product ID and lead status are still null. We need unstructured replies turned into structured lead data.
- Validation for email addresses and phone numbers.
- Human in the loop: notifying a person and asking for an action.
- Multi-agent systems come only after single-agent memory, state and resumption are solid.
