FervorCreative AI
Live Latest 02.10.26 · morning 98 tools tracked 331 workflows indexed 260 topics Hot: MiniMax H3, ComfyUI, Ming-Image-0.1-Design

Comfy Agent removes the node-graph wall by building ComfyUI workflows from a plain description while you watch, but the useful skill is directing it like an assistant (say what to keep, approve each run, save your house rules as skills), because its tokens and your renders come out of the same credit pool.

Comfy AgentComfyUIComfy Cloudcreative-workflowslocal-creative-aiprompt-craft

Comfy Agent: How to Build a ComfyUI Workflow by Describing It (and Why to Leave It on Ask)

ComfyUI's new in-app assistant builds the node graph for you while you watch. The trick is directing it like a junior assistant, not a vending machine.

ComfyUI has always had a strange reputation. It runs almost every open image, video and audio model worth using, usually first, and it scares off most of the people who would get the most out of it. The reason is the screen: a canvas of boxes and wires, where a single image-to-video setup can mean a dozen nodes you have never heard of, wired in an order nobody explains.

On October 1, Comfy shipped an answer that would have sounded like a joke two years ago. You type what you want into a chat panel, and an assistant called Comfy Agent builds the graph on the canvas in front of you. It finds the template, adds the nodes, wires them, checks that nothing is missing, and then asks before it runs anything.

That changes who ComfyUI is for. It also creates a new way to waste money, because the assistant's thinking and your renders are paid from the same account. Here is how to use it well on your first afternoon.

What Comfy Agent actually does

Comfy Agent lives inside the ComfyUI interface itself, not in a separate app. According to Comfy's documentation, it can find templates, models and nodes already available in your setup, build a workflow from a template and change its prompts and settings, add, connect and remove nodes, and check a workflow before running it. It can also look at images you attach and read the length, frame rate and format of video and audio files.

Two behaviors matter more than the feature list.

It starts from templates. Comfy's docs say that for a new workflow, the agent searches for a matching template before building from individual nodes. That is the right instinct. Comfy's templates are tested starting points, so the agent spends its effort adapting something that already works instead of inventing a graph from scratch.

You work on the same canvas at the same time. Comfy's announcement stresses that you and the agent can edit the same canvas without interrupting each other. You can drag a slider while it wires up an upscaler. That makes it less like ordering from a menu and more like sitting next to an assistant who knows the software better than you do.

Comfy calls it "the first agent for craft," and the framing is fair. It is not making the picture. It is doing the plumbing so you can spend your time on the picture.

Where it runs, and what it costs

Today, Comfy Agent runs in Comfy Cloud, the browser version of ComfyUI on Comfy's own graphics cards. The docs list it as generally available there. A version for Comfy Desktop, the app you install on your own computer, is "coming soon," with a waitlist. Comfy's blog post calls the agent a beta and says desktop is weeks away.

The cost is the part to understand before you start. Comfy gives new users some free tokens to try a few chats. After that, the agent charges "from the same Comfy Credits pool you already use for generations." The model behind it is Anthropic's Opus 4.8, and Comfy publishes the rates per thousand tokens (a token is roughly a short word's worth of text): about 1.06 credits for text the agent reads, 5.28 credits for text it writes, and 0.11 credits for text it has read recently and cached.

To put that in creative terms, here is my own rough arithmetic, not Comfy's. Comfy Cloud's Standard plan works out to $192 a year for 50,400 credits, about 262 credits per dollar on the yearly price. Comfy's pricing page says a five-second video from its Wan 2.2 image-to-video template at default settings (640 by 640, 81 frames) uses roughly 11 credits, about four cents. If one agent turn reads 30,000 tokens of workflow and conversation without the cache helping, that is around 32 credits: the price of roughly three of those five-second clips, just for the assistant to think. (A "sampler" in that example, by the way, is the node that does the actual drawing, step by step.)

That is not expensive. It is also not free, and it is invisible unless you watch. Which brings us to the most important switch in the panel.

Leave it on Ask

Next to the message box is a setting called Run permissions, with two choices. Ask makes the agent request your approval before each workflow run. Auto lets it run workflows without asking.

Auto is tempting. You describe a look, tell the agent to keep iterating until it is right, and walk away. The problem is that "right" is your judgment, not the agent's. An assistant on Auto can run ten generations chasing a result you would have rejected after the second, and both its thinking and those renders draw down the same balance.

Ask keeps you as the director. The agent builds, validates and then stops, and you decide whether this version is worth a render. For your first weeks with it, that pause is also where you learn. You see what it built before it spends anything, and you start recognizing the nodes.

Save Auto for jobs where the agent's judgment is enough, like batch-processing the same workflow across a folder of images you have already approved the look for.

Put this into practice

Here is a first session that will teach you how the agent thinks, on a job most creatives actually need: turning a still image into a short moving shot.

1. Open a blank workflow. In Comfy Cloud, open a new blank workflow, then click Ask Comfy Agent in the workflow bar and accept the consent notice if it appears.

2. Point it at the right tab. In the Agent panel, use Choose a workflow and select your blank workflow. Do this every time. Comfy's docs are explicit that "The Agent cannot create or switch tabs itself," and the most common confusion is the agent editing a different workflow from the one you are looking at.

3. Set Run permissions to Ask. Do it before you type anything.

4. Describe the shot, not the nodes. Attach your still image with Attach a file or by dragging it into the message box. Then write something like: "Turn this into a five-second clip with a slow push-in toward her face. Keep the colors and lighting from the photo. Use whatever image-to-video template fits." You are describing the result and naming what to protect. Leave the plumbing to the agent.

5. Read what it built. Before approving, look at the canvas. Ask it questions: "Why did you pick this model?" "Which node controls the length?" Each answer is a free lesson in how ComfyUI fits together.

6. Approve one run, then change one thing. After the first render, give a targeted note, and say what to keep. Comfy's docs show the pattern: tell it to "Keep the model and sampler settings" while you change something else. "Keep everything, but make the push-in slower" gets a small edit. "Make it better" gets a rebuild.

7. Save your rules as a skill. Once you have notes you keep repeating ("always 16:9," "always add a final upscale," "never change my prompt wording"), save them as a skill. Comfy's docs describe skills as plain-language instructions the agent remembers for future requests, and Comfy publishes public skills you can ask it to use. A skill is your house style, written once.

By the end, you will have a working image-to-video workflow you can reuse without the agent, and a better map of ComfyUI than an afternoon of tutorials would give you.

Where it falls short

It is cloud only for now. If the reason you use ComfyUI is that your work stays on your own machine, with your own models and custom nodes, wait for the desktop version. Comfy's agent page promises "Your GPU, your models, your custom nodes" for the local version, but it is not out.

It is a beta. Comfy's blog post calls it one and says the team is still fixing issues. Expect the occasional wrong node, wrong wire or wrong guess at what you meant. Checking the canvas before approving is how you catch it.

It costs credits twice. The agent's tokens and your renders draw from one pool. A long conversation about a big workflow is the expensive case, so start a fresh chat for a new job rather than dragging one conversation through an entire afternoon. Comfy's agent page says you can run up to five chats in parallel, each with its own history.

It cannot manage your tabs. You select the workflow; it edits that one. Forget, and it edits the wrong graph.

It is only as good as your description. The agent can wire nodes. It cannot know that your client hates lens flare. The quality of your notes is still the quality of the work.

The plumbing was never the craft

For years, ComfyUI has sorted creative people into two groups: those willing to learn node graphs and those who were not. A lot of excellent illustrators, editors and motion designers ended up in the second group, not for lack of taste, but for lack of patience with wires.

Comfy Agent does not remove the need to know what you want. It removes the tax you paid to express it. The people who will get the most from it are the ones who already direct: who know what to keep, what to change, and when a take is good enough to stop.

So try one shot this week. Describe it, keep the agent on Ask, and watch what it builds. You might find that the scariest screen in creative AI was only ever waiting for a clear brief.


Medium metadata

Title: Comfy Agent: How to Build a ComfyUI Workflow by Describing It (and Why to Leave It on Ask)

Subtitle: ComfyUI's new in-app assistant builds node graphs from plain English while you watch. Here is a first-session walkthrough, what it costs in credits, and where it still trips up.

Tags: ComfyUI, AI Video, Generative AI, Creative Workflow, AI Tools

Estimated read time: 8 minutes