Picking an agent framework: PydanticAI, LangChain or Strands
I get asked which agent framework to use about once a week. I use PydanticAI. This post is how I got there, with a fair look at the two I get asked about most as alternatives: LangChain, and Strands from AWS. I've added Marvin too, because it turns up the minute you start reading about Prefect.
I get asked which agent framework to use about once a week. I use PydanticAI. This post is how I got there, with a fair look at the two I get asked about most as alternatives: LangChain, and Strands from AWS. I've added Marvin too, because it turns up the minute you start reading about Prefect.
#What an agent framework is for
Stripped down, an agent is a loop. You send a model a task and a list of tools. The model asks for a tool. Your code runs it and sends back the result. That repeats until the model says it's done.
You can write that loop yourself. My flower shop demo does it by hand in about forty lines, straight against a model SDK. I did that on purpose so people could read every line.
Those forty lines grow, though. You want the answer back as a real object and not a string you have to parse. You want to change models without a rewrite. You want tool arguments checked before anything runs, a way to stop for approval, tests that don't call a paid API, and traces. A framework is somebody else's well-tested version of all of that.
These are the questions I ask of one:
- How much does it hide? When something goes wrong, I need to see what was sent to the model.
- Is the output typed? I want a validated object back, with fields my editor knows about.
- Can I change models? And does it quietly tie me to one cloud?
- How does a risky tool get stopped? Sending a quote or spending money needs a person's yes first.
- Can I test without paying for model calls? Tests have to run in CI with no API key.
- Does it get along with the rest of my stack? For me that's Prefect for orchestration and OpenTelemetry for traces.
- How often does it change underneath me?
#PydanticAI
PydanticAI comes from the people who make Pydantic, and it feels like it. Everything is typed. You say what the agent returns as a Pydantic model and you get that object back, already validated:
from pydantic import BaseModel
from pydantic_ai import Agent
class Quote(BaseModel):
customer: str
amount: float
summary: str
agent = Agent(model, output_type=Quote)
result = agent.run_sync("Quote the June wedding inquiry")
print(result.output.amount)
The model is a string you pass in, and it supports nearly every provider, so nothing about it pins me to a vendor.
The approval story matches how I already build. You mark a tool requires_approval=True. Or, when it depends on the arguments, the tool raises ApprovalRequired itself, say for an order over a certain amount. When the model calls that tool, the run ends and hands you the pending calls. Later, in another process if you like, you start a new run with the decisions attached. Nothing is held open in memory and there's no trick to it. The docs also say outright that approval doesn't replace authorization checks inside the tool, which is the point I keep making to clients.
It ships a test model, so my tests run in CI with no API key. Tracing is OpenTelemetry from the start. And it has documented, co-maintained integrations with Prefect and Temporal for runs that need to survive a crash.
My complaints are about pace and size. It moves fast. The way you connect it to Prefect changed inside a year, and the old way is already deprecated. The community is also smaller than LangChain's, so there are fewer ready-made integrations and fewer answered questions when you get stuck.
#LangChain
LangChain is the one everybody has heard of, and for a long time it had a reputation for too many layers between you and the model. Version 1.0 was a real cleanup. There's one way to make an agent now, and the older chain machinery moved out to a separate package.
from langchain.agents import create_agent
agent = create_agent(
model="claude-sonnet-4-6",
tools=[search_web, analyze_data, send_email],
system_prompt="You are a helpful research assistant.",
)
result = agent.invoke(
{"messages": [{"role": "user", "content": "Research AI safety trends"}]}
)
Under that sits LangGraph, which is the runtime. It saves state, survives crashes and can pause for a person. So with LangChain you get the agent framework and the orchestration piece from the same people.
Approvals are handled by middleware. You list which tools need review, and the reviewer can approve, edit or reject each call. Being able to edit the arguments before a call runs is a nice touch. It needs a checkpointer to store the paused state, and with a Postgres one the pause survives a restart. There's also built-in middleware for redacting personal data and for trimming long conversations.
And the ecosystem is the biggest by a wide margin. If you need a connector to something obscure, LangChain probably has it.
What it costs you is surface area. There's LangChain, LangGraph, LangSmith for tracing, middleware and checkpointers, and you need a working picture of all of them. LangGraph also wants to be the runtime, and I've already picked an orchestrator. I don't want two tools that both think they own retries and waiting. You also spend more time in message dictionaries and less in typed objects than you do in PydanticAI. And anybody who used it before 1.0 remembers how often the API moved. Version 1 promises stability, and that's a promise it still has to keep.
#Strands
Strands is the open-source agent SDK from AWS, in Python and TypeScript. They call the approach model-driven. You hand it a model and some tools and let the model work out the plan. It takes the least code of the three:
from strands import Agent, tool
@tool
def letter_counter(word: str, letter: str) -> int:
"""Count occurrences of a specific letter in a word."""
return word.lower().count(letter.lower())
agent = Agent(tools=[letter_counter])
agent("How many R's are in strawberry?")
Notice there's no model named in that example. The default is a Claude model on Amazon Bedrock, so if your AWS credentials are set up, it just runs. You can point it at other providers too, including Anthropic directly, OpenAI and Ollama.
It has a lot built in: MCP support, OpenTelemetry tracing, and several multi-agent patterns such as agents as tools, swarms and graphs. For approvals it has interrupts. A hook that fires before a tool call can stop the agent and ask a person, or a tool can raise the interrupt itself. A session manager saves the paused state so it survives a restart.
Strands is at its best on AWS. It deploys straight to AgentCore, where you also get AWS's identity service and policy engine for agents. If a client already lives in AWS and their security team has already approved Bedrock, this is the shortest road to something running.
The pull toward AWS is also my main complaint. The defaults assume Bedrock, the best deployment story is AgentCore, and it's easy to end up more committed to one cloud than you meant to be. It's the youngest of the three. And model-driven means less of the control flow is written down in your code, which is pleasant right up until you have to explain exactly why the agent did something.
#Marvin
Marvin comes from the team behind Prefect, which is how it landed on my list. It's a different kind of thing from the other three. Marvin 3 is built on top of PydanticAI, so it's a layer over a framework more than a framework in its own right. (It also replaced ControlFlow, Prefect's earlier agent project, which is archived now.)
What it adds is convenience. The simplest call is one line:
import marvin
poem = marvin.run("Write a short poem about artificial intelligence")
print(poem)
It thinks in tasks. A task is a clear objective with a typed result, handed to an agent. Threads carry the conversation history from one task to the next. And it has a handful of one-line helpers (extract, cast, classify and generate) for turning messy text into structured data. Those are handy for small jobs.
Where it falls short for me is approvals. A task can ask a person for input at the command line by setting cli=True. I didn't find a way to hold one specific tool call for sign-off, or to keep a pause alive from one process to the next. PydanticAI underneath does both, so by going through Marvin I'd give up the thing I care about most. It's also been quiet. The latest release on PyPI is from March.
I'd reach for Marvin for a quick structured-output script. For an agent that can spend money, I use PydanticAI directly.
#The heavy hitters I skipped
There are five more names people expect in a post like this: the OpenAI Agents SDK, Google's Agent Development Kit, Microsoft Agent Framework, the Claude Agent SDK and CrewAI. I haven't put any of them through their paces. I'd rather tell you why than pretend I ran a bake-off.
The Claude Agent SDK is opinionated about the model. It's Claude Code packaged as a library, with the Claude Code binary bundled in, so it runs Claude and only Claude. You can reach Claude through Anthropic or through a cloud provider that serves it, but that's the extent of the choice. It also arrives with file and shell tools built in, because it's a coding agent at heart. My demo runs on Claude today. I still don't want my framework making that decision for every client.
I reach for AWS because it's where I deploy, and that's why Strands got a full section above. Google's kit and Microsoft's framework are the same idea from the other two big clouds, and I'm not deploying to either one. A client on Google Cloud or Azure wouldn't change that. I've already picked a framework that doesn't care which cloud it runs on, so PydanticAI comes with me.
The OpenAI Agents SDK can run other providers' models, and its own docs list what you give up when you do: the hosted tools, some structured output support, and tracing that reports to OpenAI unless you switch it off. I'm not using OpenAI models, so I'd be working against the defaults from the first line.
CrewAI I wrote off early. The buzz around it read as spam to me, and I never looked that direction again. What I've read since hasn't pulled me back: security advisories this year, reports of runaway token costs, multi-agent runs that are hard to debug, and usage telemetry that's on by default. All of that kept me from digging deeper. If you run CrewAI in production and think I've made the wrong call, tell me, and I'll take another look.
Some of these reasons are small, and I know it. This is how tool choices get made when you're building and shipping fast. I reach for what gets recommended to me in the hallway track, I compare it against whatever has caught my own interest, and I get back to work. I don't have time to evaluate every framework side by side.
#Side by side
| Question | PydanticAI | LangChain | Strands |
|---|---|---|---|
| Who makes it | The Pydantic team | LangChain | AWS |
| Languages | Python | Python and JavaScript | Python and TypeScript |
| Typed output | The core of the design | Supported in the agent loop | Supported |
| Models | Nearly any, chosen by a string | Nearly any | Many, with Bedrock as the default |
| Stopping a risky tool | Mark the tool; the run ends and resumes later with decisions | Middleware with approve, edit or reject; needs a checkpointer | Interrupts from a hook or a tool; a session manager saves state |
| Runtime | Bring your own (Prefect, Temporal and others) | LangGraph, included | AgentCore, Lambda and other AWS targets |
| Tracing | OpenTelemetry | LangSmith | OpenTelemetry |
| Ecosystem | Smaller, growing fast | The largest | The newest, centered on AWS |
Marvin isn't in the table because it sits on top of PydanticAI. Its answers to most of these questions are PydanticAI's answers.
#What I use
My final decision came down to two options: stay cloud-agnostic, or go all in on Strands with AWS. I went cloud-agnostic, and for me that means PydanticAI. Typed output is how I already write Python. Its approval model lines up with the approval gate I build into every system. It doesn't bring its own runtime, so Prefect gets to do the orchestrating without a turf war. Traces are OpenTelemetry, and my tests run without a model. And it runs the same on AWS as it would anywhere else, so a client's cloud never reopens the question.
I'd pick LangChain when a project needs an integration only LangChain has, when the team I'm joining already knows it well, or when a client wants the framework and the runtime from one vendor with one support contract.
Whichever one you choose, keep the approval rule and the tool allow-list in your own code. All three of these will happily stop for a human. None of them knows which of your tools can cost somebody money.
If this post helped, you can support my open source work on Ko-fi.