Save time, make money and get customers with FREE AI! CLICK HERE →

DeepSeek Harness vs Claude Code: 57x Cheaper, Tested


Here is the deepseek harness vs claude code number that should change how you budget for AI: five cents.

That is what a complete 3D animated website plus a working game cost to build on DeepSeek Harness.

Kasra Dash topped his account up to $10 right before we recorded, and the build barely touched the balance.

It also finished in 11 minutes while Claude Code was still working past 30.

And yet I am still using Claude Code for client work, which is the part worth explaining.

The cost verdict first

If your AI bill is what stops you shipping more, DeepSeek Harness fixes that today.

If your bottleneck is output quality, it does not.

Both of those came out of the same test, on the same prompt, on the same afternoon.

How the test was set up

One prompt, no context, no follow-ups.

We asked for a 3D animated accountancy website and a simple game.

DeepSeek Harness ran DeepSeek V4 Pro.

Claude Code ran Claude Opus 5 on high.

That is a like-for-like frontier comparison rather than a cheap model against an expensive one.

The numbers

DeepSeek Harness finished the whole build in 11 minutes.

Claude Code was past 30 minutes and still running.

DeepSeek burned 483,000 tokens getting there.

Claude was at 48,000 tokens at the 20-minute mark.

So DeepSeek used roughly ten times more tokens, on a model priced at about a 57th per token.

Do that arithmetic and the cheap engine still wins comfortably, even while being wildly wasteful.

That is the part people miss when they quote the per-token price alone.

Cost factor DeepSeek Harness Claude Code
Tokens on this build 483,000 48,000 at 20 minutes
Relative price per token ~57x cheaper Baseline
Real cost of the run About 5 cents Materially higher
Time to finish 11 minutes 30+ minutes
Harness licence Free, open source Bundled with the product
Model choice Swappable, free brains work Claude only

Where the five cents does not help

Open both builds side by side and Claude’s is the one you would show a client.

Its animation followed the mouse rather than just decorating the page.

It also worked in details nobody typed, including Northwest England, because it had context from earlier sessions.

The DeepSeek build came out animated and cartoony, and its Tetris was functional but tonally wrong for an accountancy firm.

Kasra’s line was that those two things should never mix, and he is right.

Neither build was ready to publish.

Both needed more prompting and more context.

The difference was how far from ready each one landed.

Old way vs new way

Old way New way
One agent, one bill, whatever it costs. Several engines, each job routed to the cheapest that can do it.
Drafts burn premium tokens. Drafts cost around 5 cents.
Ration attempts to control spend. Attempts are effectively free, so you iterate more.
The vendor chooses your model. Everything is a plugin, so you choose the brain.

🔥 Want the routing setup that cuts the AI bill?

I walk through the whole Agent OS build inside the AI Profit Boardroom — cheap engine for volume, premium engine for client work, 4,000+ members and weekly coaching calls.

→ Get access here

Free harness, free brain, near-zero bill

The harness is free and open source, which is why it collected 105,000 GitHub stars in about two days.

That makes it one of the fastest growing open source projects anyone has tracked.

But a harness on its own does nothing.

It needs a brain plugged into it.

You can pay for DeepSeek V4 Pro, which is what we tested, and the bill still rounds to pennies.

Or you can plug a free model in instead, and OpenCode works as a free brain inside the harness, which takes the running cost to zero.

When an attempt costs nothing, you stop rationing attempts, and that changes how you work more than any single feature does.

A worked example

Say you generate twenty scaffolds a week: landing pages, component variations, first drafts.

At around 5 cents and 11 minutes each, that is roughly a dollar a week and about four hours of wall-clock time.

The same workload on a premium engine runs at roughly 57 times the price per token, on an agent that took over 30 minutes for the same brief.

The arithmetic is not subtle.

The trap is concluding that the cheap engine should therefore do everything.

Our test showed the cheap output landing further from finished, and every extra prompting round costs you time even when it does not cost you money.

The rule I actually use

If the output goes in front of somebody who pays me, it runs on the premium engine.

If the output is an input to my own next step, it runs on the cheap one.

That one rule captures most of the saving without the risk.

And I do not switch by hand.

My agent operating system acts as the orchestrator and routes each job to the right engine.

I never even opened the harness to install it — I asked Claude to set it up, test it and wire it in.

The risk sitting inside the price

DeepSeek Harness is a v0.1 developer preview, and its pricing is early-stage pricing.

Nothing stops it rising, and a very verbose model at a normal price is a far less exciting product.

So build your workflow so the engine is swappable.

That is the entire argument for a harness where every part is a plugin, and it is the argument for putting an orchestrator in front of your tools rather than marrying one of them.

It is also why I gave the harness a 7 out of 10 while Kasra scored it lower — he was rating today, I was rating the trajectory.

What the test does not tell you

One prompt is one data point.

If your work is more repetitive than ours, the cheap engine looks even better.

If it leans on accumulated context, the premium engine pulls further ahead than our numbers suggest.

Run the same test yourself: one prompt, both agents, no extra context, and watch the token counters as well as the output.

On the DeepSeek side it costs about 5 cents to find out.

Why 105,000 stars is a vote about ownership

The GitHub number is easy to misread as a quality score.

It is not one.

DeepSeek Harness did not out-build Claude Code in our test, and the stars still went vertical.

What people are voting for is control.

The harness runs on your machine, the model is your choice, and every part of it can be pulled out and replaced.

For the last year, if you wanted a serious coding agent you had one obvious answer and you paid whatever it cost.

Now there is a free, open, fast alternative with a frontier model behind it and a community shipping plugins for it within days.

That does not have to beat Claude to matter.

It only has to be good enough that Claude cannot stand still.

Which is the real reason to spend an afternoon on it even if you are staying where you are — you are not evaluating a replacement, you are building the option to leave.

That option is the thing that keeps everybody’s pricing honest.

Speed is really a measure of attempts

Eleven minutes reads like a bragging stat until you convert it.

An agent that finishes in 11 minutes gives you roughly five attempts in an hour.

An agent that takes over 30 gives you one, maybe two.

Almost nothing good comes out of the first attempt, so the tool that lets you fail four more times per hour has a real advantage — provided failing is cheap.

At around 5 cents a build, failing is effectively free.

Fast plus cheap is what makes this interesting even though Claude produced the better page.

You are not comparing one polished output against one polished output.

You are comparing one polished attempt against five rough ones, and for a lot of internal work five rough ones is the better trade.

For client work it is not, which is exactly why both stay installed on my machine.

Where to start this week

Install it through an agent you already use. Do not learn a new interface to save money; get your existing agent to stand it up and wire it in.

Point it at a free brain first. Pay for the premium model only when a job justifies it.

Split your recurring jobs into client-facing and internal. That list is your routing table.

Measure token burn on a real job. Toy prompts hide verbosity; 483,000 tokens on one page is what it looks like at scale.

Keep the premium engine wired in. Our test showed exactly where the cheap one falls short, and it falls short in the place clients look first.

Frequently asked questions

How much cheaper is DeepSeek Harness?

Roughly 57 times per token. Our build cost about 5 cents, though it used around ten times more tokens than Claude for the same job.

How many tokens did it use?

483,000 for a single page, against Claude’s 48,000 at the 20-minute mark.

Is the harness free?

Yes, free and open source. You supply the model, and free brains work inside it.

Which built the better site?

Claude Code, clearly — mouse-reactive animation and business context it was never given.

How should I run both?

Route per task with an orchestrator: volume to DeepSeek, client work to Claude.

About Julian

I’m Julian Goldie, founder of a 7-figure SEO and link building agency (Goldie Agency, 70+ team) and the AI Profit Boardroom.

I have 400K+ YouTube subscribers, 163K X followers and 29K+ Udemy students, and I wrote Link Building Mastery.

Everything above came from a filmed side-by-side test, not from a spec sheet.

Also on our network

Different angles on the same test: juliangoldie.com and goldstarlinks.com.

📺 Video notes + links to the tools 👉

🎥 Learn how I make these videos 👉

🆓 Get a FREE AI Course + Community + 1,000 AI Agents 👉

The deepseek harness vs claude code cost gap is real — spend it on volume, not on your best work.