The Bajoran Engineer

Why I'm moving my agents from Metaflow to Prefect

I'm moving my agent orchestration off Metaflow and onto Prefect. If you read the notebook post last week, you already watched Prefect do the scheduling, so this probably isn't a huge shock.

Why I'm moving my agents from Metaflow to Prefect

I'm moving my agent orchestration off Metaflow and onto Prefect. If you read the notebook post last week, you already watched Prefect do the scheduling, so this probably isn't a huge shock.

This isn't a Metaflow hit piece. I like Metaflow! But it got in my way in a few specific places, and those places are the useful part, so that's where I'm spending my time.

Full disclosure: I walked into this already leaning toward PydanticAI for the agents themselves, and that tilted the whole decision. If you want the why on PydanticAI, I wrote a whole post about it.

#What does an orchestrator even do for an agent?

The agent framework defines the agent: its instructions, its tools, and the loop where it decides what to do next. Mine is PydanticAI.

The orchestrator is everything around that. When does a run kick off? What happens when a step falls over? Where does the record of all of it live? What happens when a run has to stop and wait on a human?

You can skip it for a while. A script and a cron job will get a demo out the door, no problem. Then a run dies after four model calls you already paid for, or a quote needs the owner's sign-off, or a client asks what happened last Tuesday... and surprise, you've been writing an orchestrator by hand this whole time.

So this is my checklist now:

  • How does a run start? On a schedule, when something happens (an email comes in, a new order lands), or because I asked it to.
  • What happens when a step fails? I want to retry that step without redoing (and paying for!) the ones that already worked.
  • Can a run wait for a person? For two minutes or most of a day, and then pick up right where it stopped.
  • Can someone who isn't me see what ran? The shop owner shouldn't need a terminal.
  • What happens when the job gets big? Like tens of thousands of documents, not ten.
  • Does it know about my agent framework? I want every model call and tool call tracked on its own, not one big black box.
  • What does it cost to keep running? How many moving pieces, and who has to babysit them?

#Why now?

I'm locking in the stack I use for client agent work, and the orchestrator is the piece everything else hangs off of. My evals run through it. My traces come out of it. My approvals sit in it, waiting. I'd rather pick it once, on purpose, and write down why, than end up with whatever I happened to grab first.

Metaflow is what I'd been using. When I held it up against that list, it absolutely crushed the big-job question and came up short on most of the rest. Prefect is good enough on all seven for the kind of work I do. So, decision made. The rest of this post is me showing my work.

#What Metaflow does really well

Metaflow is built for flows that do a LOT of computing. You write your steps as methods on a Python class, run it on your laptop, and slap a decorator on any step that should run in the cloud. Every run saves what each step produced, and resume restarts a failed run at the step that broke. Version 2.18 added conditional and recursive steps, so a flow can loop on whatever a model decides. Their write-up shows an agent chewing through hundreds of documents in parallel, and it's pretty slick.

That's a real strength, and big jobs are part of what I build. So before I moved anything, I had to make sure I wasn't giving that up. (More on that below.)

#Agents need to be able to wait

My agents do real work. They read orders, check stock, draft quotes and put together the morning brief. Some of that takes a few seconds and some of it takes a while.

And at the important moments, they have to hold. A quote gets drafted, and then nothing should happen until a person says yes. Maybe that takes two minutes. Maybe it takes nine hours, because the owner is on the shop floor all day and is not checking their phone. (That's the flower shop demo, and my client builds work the same way.)

So I want an orchestrator that's chill about all of that. Run now. Run on a schedule. Run when an email lands. Stop in the middle and pick back up tomorrow. I don't want to design around my tools every single time an agent needs to wait on somebody.

#Where Metaflow got in my way

It can't schedule itself. Locally a flow is one command, but the docs are upfront that production means handing it to another orchestrator: Argo, Step Functions, Airflow or Kubeflow. So my orchestrator needs... an orchestrator. For a morning brief at a three-person shop, that's a whole lot of machinery sitting behind what's basically a cron job.

There's a lot to stand up. A production setup wants object storage, a metadata service, a compute backend and that second scheduler. Somebody has to own the Terraform and the IAM for all of it, and at a small client, that somebody is me.

No dashboard out of the box. The UI is a separate thing you deploy on top of the metadata service. Day to day, it's the command line and notebooks. Which works fine for me! It does nothing for a shop owner who wants to see what ran this morning.

It still doesn't run natively on Windows. You need WSL, and a lot of small businesses are Windows shops.

It leans hard toward AWS. It can run elsewhere, but the easy path and most of the examples assume AWS. I'm not the only one who's noticed; it comes up in most of the comparisons I read.

There's no way to stop and wait for a person. I looked! The agent material is all about autonomous runs, and my whole design is about the moment a run stops being autonomous, which is exactly when money is about to move.

It doesn't know my agent framework exists. I write agents in PydanticAI. You can absolutely call a PydanticAI agent from inside a Metaflow step, it's all Python. But nobody documents that pairing, and Metaflow sees the whole agent run as one opaque step. Prefect has an integration that the Pydantic team co-maintains, so every model request and tool call gets tracked and retried on its own. One of these tools is keeping up with my framework, and the other one leaves the wiring to me.

None of these alone would've made me leave. But stack them all up and I'm building the scheduling, the waiting, the visibility and the glue myself. No thanks.

#So, Prefect

Quick clarification, because I had to sort this out for myself: Prefect is an orchestrator, not an agent framework. It doesn't care how the agent thinks. It runs things on a schedule, retries them, writes down what happened, and knows how to wait.

Hooking it up to PydanticAI is one line on the agent:

python
from prefect import flow
from pydantic_ai import Agent
from pydantic_ai.durable_exec.prefect import PrefectDurability

agent = Agent(model, name="orders", capabilities=[PrefectDurability()])

@flow
async def handle_inquiry(task: str) -> str:
    result = await agent.run(task)
    return result.output

Now every model request and tool call is its own tracked task. If a run dies halfway through, the retry skips whatever already worked, so I'm not paying twice for the same model calls. That alone would've gotten my attention.

It can also stop and ask a person something:

python
from prefect import flow, pause_flow_run
from pydantic import BaseModel

class Approval(BaseModel):
    approve: bool
    note: str = ""

@flow
def quote_flow():
    draft = draft_quote()
    decision = pause_flow_run(wait_for_input=Approval)
    if decision.approve:
        send_quote(draft)

Prefect builds a form from that model and hands me back a type-checked answer. There's also a suspend version that shuts the process down while it waits.

One thing I want to be really clear about: that pause is a waiting room. That's it. My approval gate is its own separate thing. The rule about what needs a yes, and the check that nothing goes out without one, stay in my code where I can write tests against them. I'm not handing that off to an orchestrator, no matter how nice the form looks.

And it's small! Self-hosted, it's two Python processes and a database: the Prefect server, which keeps the schedules and serves the dashboard, and a worker that picks up the runs. Both fit on a little hosted service next to Postgres, or I can use Prefect Cloud and skip running the server. Either way there's a dashboard, so the owner can see what ran without asking me. When your client is a shop with three employees, "cheap to leave running" beats most features.

#Okay, but what about the big jobs?

This was my real worry. Running an agent over tens of thousands of documents is exactly where Metaflow shines, and I do that kind of work. So can Prefect hang?

It can! There are three ways in, depending on how big the job is.

The simplest is fanning out inside one flow. .map() runs the same task over a list of inputs, and by default that happens in threads. Since an agent task spends most of its time waiting on a model API anyway, threads get you pretty far.

When one machine isn't enough, you swap the task runner. Prefect has runners for Dask and Ray, and both can spread tasks across a cluster:

python
from prefect import flow, task
from prefect_dask import DaskTaskRunner

@task
def review(doc):
    ...

@flow(task_runner=DaskTaskRunner())
def review_all(docs):
    return review.map(docs)

And when the whole job needs bigger hardware, you send it there. Work pools let one deployment run on a small box and another on Kubernetes, ECS or Cloud Run. There are also decorators that ship a single flow off to that infrastructure straight from Python.

Here's the catch, though. Prefect coordinates the compute, but you have to bring it. I'm the one supplying the Dask cluster or the Kubernetes work pool, and those decorators need a work pool and object storage set up first. In Metaflow, asking for a bigger machine on one step is a single line. So Metaflow is still easier for the giant job. Prefect is good enough for mine, and I get to keep one tool for everything.

#"Just use Temporal"

I know, I know. Ask around and that's the answer for serious agent work, and PydanticAI supports it too.

I think Temporal is the right call for some systems; it's just more than I need right now. It wants a cluster or a paid service. Workflow code has to be deterministic. Anything that crosses into a tool call has to serialize, and there's a payload size limit to plan around. What you get in return is a workflow that can wait a week for a signal and wake up exactly where it left off, which is genuinely amazing.

But the approvals I'm building for come back the same day. I'd be signing up for all of Temporal's rules to get guarantees I mostly don't need yet. The day a client's approvals sit for a week, or a client already runs Temporal, I'll use it and be glad it exists.

#What I'm giving up

The artifact store. That one hurts. Every run, every step, every output, saved and versioned without me lifting a finger. Prefect keeps states and results, but that's not the same as opening an old run in a notebook and poking around. I'm counting on traces to cover the gap, and I don't know yet if they will. Ask me in a few months.

I'm also giving up the easy version of scale. Like I said above, big jobs on Prefect mean I set up the cluster, which means more of my time on every large job.

And the pause has a catch. Prefect has two ways to wait. A pause blocks: the flow's process keeps running until someone answers, and by default it gives up after an hour. A suspend exits the flow and frees up the machine. For a quote drafted at 7am and approved at 4pm, suspend is the one I want.

But suspend has its own cost. When the answer comes in, Prefect reschedules the flow and it runs again from the top. It only skips the finished steps if their results were saved. So the flow has to run as a deployment with result persistence turned on, or the agent redoes its work and I pay for the same model calls twice. (My short example above doesn't show any of that. Consider this your footnote.)

I'm still rooting for Metaflow, for real. Its one-line scaling is still the easiest around. But as of today I haven't seen human-in-the-loop approval or a documented PydanticAI integration, and I need both. When it has them, I'll take another look.

If you've run agents on either one and think I'm missing something, tell me! I'd love to hear it.

If this post helped, you can support my open source work on Ko-fi.