Summary of Today’s Gen AI Project: LangGraph Memory, Threads and Summarization (3 Oct 2026)

Today’s project session was one long debugging arc: prove the agent has no memory, give it memory, then stop that memory from bankrupting us.

1. Proving the Problem

We compiled the existing graph as graph_v0 and rendered it as a Mermaid PNG with IPython.display to confirm the flow visually.

Then the live test:

  • Turn 1: asked for a recommendation → the agent suggested the Radiancy Vitamin Cream.
  • Turn 2: “which product did you just recommend?” → the agent denied recommending anything.

Printing the message history explained it: turn 2 didn’t contain turn 1’s messages. Without a checkpointer, every .invoke() is a brand-new session.

2. Checkpoints and Thread IDs

from langgraph.checkpoint.memory import MemorySaver

checkpointer = MemorySaver()
graph_v1 = state_graph.compile(checkpointer=checkpointer)

cfg = {"configurable": {"thread_id": "customer-1"}}
graph_v1.invoke({"messages": [...]}, cfg)

Two things, both required:

  • The checkpointer saves state after each step. Note the keyword is checkpointer, not checkpoint.
  • The thread ID says whose conversation this is. It must sit inside configurable.

Proof it worked: turn 2 went from 2 messages to 6. The history was being carried.

3. Inspecting State

snapshot = graph_v1.get_state(cfg)
snapshot.config        # contains the checkpoint_id
snapshot.values        # iterate to print each message

MemorySaver keeps all this in RAM. In production it belongs in a database.

4. Simulating Multiple Customers

We wrapped it in a chat function that reads input, exits on exit, quit or an empty line, and passes a thread ID through.

  • Customer 1: a skin concern, quoted a serum at ₹699 for 30ml
  • Customer 2: oily skin, recommended a cleanser at ₹349 for 100ml

Separate thread IDs kept the two conversations completely separate, and resuming an old thread ID picked the conversation back up.

5. Two Flaws We Found

Nodes returning the whole state. Inefficient, and it risks overwriting the message history. Return only what changed:

def assistant(state):
    return {"messages": [reply]}      # not the entire state

Nothing is persistent. MemorySaver dies with the process. Production needs SQLite or Postgres:

# pip install langgraph-checkpoint-postgres
from langgraph.checkpoint.postgres import PostgresSaver

with PostgresSaver.from_conn_string(DB_URI) as checkpointer:
    checkpointer.setup()        # creates tables on first run
    graph = builder.compile(checkpointer=checkpointer)

6. The Cost Problem

The messages list grows with every turn, and you resend all of it each time. At 50+ turns the token bill stops being trivial.

Since it’s just a list, we can prune it. The plan: summarise once the conversation gets long. We accept slightly lower fidelity in exchange for bounded cost.

7. The Summarization Node

ParameterValueNotes
SUMMARY_TRIGGER12 messages20 was the first proposal; tune it
KEEP_RECENT6 messagescould be 50–150 on a larger-context model
Summary lengthunder 120 wordsmay need to grow to ~500 in testing

The prompt rules: keep the customer’s status, their concerns and the specific products discussed, and never invent anything.

State and continuity

  • Added a summary string field to the advisor state.
  • Each new summary extends the previous one rather than replacing it, so nothing silently drops.
  • The LLM call returns the summary; we extract just the content.

Cleaning up old messages

from langchain_core.messages import RemoveMessage

return {"messages": [RemoveMessage(id=m.id) for m in old_messages]}

A safe_cut_messages helper decides how many to drop while always preserving a human message, so the history never starts mid-thought. RemoveMessage only works when the state key uses the add_messages reducer.

Routing

  • route_after_assistant: tool calls present → the tools node; otherwise → the summarization route.
  • route_exit: message count over the trigger → the summarize node.
  • We used matching node and route names to keep the conditional mapping simple.

8. Worth Comparing: The Built-In Option

LangGraph and LangMem already ship trim_messages, summarize_messages and a prebuilt SummarizationNode. Build yours first, then read theirs and see which decisions they made differently. That comparison is the real exercise.

✅ Your Assignment

Treat today’s code as a draft from a developer, and yourself as the tester.

  1. Clone it, run it, and debug the summarization flow.
  2. Any failure is a defect for you to fix, not a question to bring to class.
  3. Fix the LLM prompt so the model actually uses the summary field. Right now it’s written but not consumed.
  4. Test with real conversation data and find the word limit that works.
  5. Convert the nodes to partial state updates.

9. Still Open

  • Structured extraction: customer name, mobile, product ID and lead status are still null. We need unstructured replies turned into structured lead data.
  • Validation for email addresses and phone numbers.
  • Human in the loop: notifying a person and asking for an action.
  • Multi-agent systems come only after single-agent memory, state and resumption are solid.

By continuous learner

enthusiastic technology learner

Leave a Reply

Discover more from Direct AI Powered By Quality Thought

Subscribe now to keep reading and get access to the full archive.

Continue reading