We ve gotten pretty good at measuring what an LLM costs.
Tokens in.Tokens out.API bill.
But an agent does a lot more than generate tokens.
It might call three tools, run a search, query a database, invoke another agent, retry a failed step, and finally produce something useful.
So what did that piece of work actually cost?
That s the problem I ve been thinking about with Kopai.
Every agent run should have a ledger:
What did it do?
What did it consume?
What did it cost?
And eventually what did it earn?
Not just token accounting.
A financial ledger for agentic work.
Tell me what you are using to monetise you Agents
Hey Omri, thanks for actually building something and being specific about what worked, that's more useful to us than a star rating alone.
On the eval gate, I appreciate you noticing that, it's been our main focus. Every agent gets scored across dimensions (system prompt quality, scope adherence, safety, plus KB/tool accuracy where relevant) before it can publish, and any edit re-triggers the check, so it's not publish-once-and-hope. The passing threshold is 70/100 and it's visible in the builder next to the per-dimension breakdown, not a black box.
About live vs frozen snapshot, it pulls the current published version, not a snapshot frozen at embed time. So when you update the agent, anyone who's embedded it sees the update too. Fair callout that this wasn't clear up front, we'll get it stated explicitly wherever embed is set up.
On technical categories being thin; Fair, and it tracks who's found us first, mostly consulting, coaching, and finance backgrounds. Getting devops/infra/security folks publishing is next, if you know anyone in that world sitting on knowledge worth packaging, send them our way!