kaedax

AI · Agents · 720-hour cycle

AI agents that survive
contact with production.

kaedax builds production AI agents and AI features — agent products, copilots, AI-powered workflows — in a fixed 30-day (720-hour) cycle. The evaluation harness ships before the agent: graded test sets, regression gates, refusal layers, and MCP tool contracts. We run our own delivery on agents, so we build what we operate.

ops.tallow.ai/evals

Eval suite

412/412

green

Task success

96.8%

+1.2 vs v12

Refusal accuracy

99.1%

0 false-do

Cost / task

$0.041

−18% mom

Eval setCasesPassΔ vs prodGate
core-tasks.v1318098.3%+0.6 pass
adversarial.v89696.9%+2.1 pass
refusals.v872100%0.0 pass
tool-contracts.v544100%0.0 pass
long-horizon.v22085.0%−3.0 blocked
release v13 gated · long-horizon regression under review · human paged
representative build · web console + mobile · anonymized under NDA
Try tallow — live agent-evals demo

[ §01 ] the problems we're hired for

Where ai & agents builds
actually go wrong.

01

The demo works; production doesn't

Every AI demo looks magical against cherry-picked inputs. Production means adversarial users, malformed data, and edge cases at volume. We build the eval harness before the agent — graded test sets and regression gates that make 'it works' a measured claim.

02

Nobody can tell you if the agent got worse

Prompt tweaks and model upgrades silently regress behavior. Without evals in CI, you find out from your users. Our builds gate every agent change behind the same green-suite discipline as code.

03

The refusal layer is the product

An agent that confidently does the wrong thing destroys trust faster than one that asks. Scoping what the agent won't do — and proving it with tests — is where enterprise buyers decide. We've written publicly on this; it closes deals.

04

Tool integrations are 80% of agent engineering

Agents are only as good as their tools' contracts. We build MCP-based tool layers with typed contracts, sandboxed execution, and audit logs of every tool call — the plumbing that separates agent products from agent demos.

[ §02 ] what ships

One cycle.
Production, not prototype.

The shapes of ai & agents work that fit a 720-hour cycle. Every build ships with runbooks, monitoring, and 30–60 days of post-launch on-call.

Eval harness first Refusal-layer tests Tool-call audit logs Cost & rate budgets Model-swap safe
[01]

Agent products end-to-end

The agent, the tool layer, the eval harness, and the operator console — shipped as one product.

[02]

AI features in existing products

Copilots, summarization, extraction — dropped into your codebase with evals gating every release.

[03]

Eval harnesses

Graded test sets, regression gates, accuracy budgets — demonstrable to enterprise buyers and auditors.

[04]

MCP tool layers

Typed tool contracts, sandboxed execution, per-call audit logs, rate and cost budgets.

[ §04 ] questions

AI & Agents founders
ask us first.

01

How do you make sure an AI agent actually works?

+

We build the evaluation harness before the agent: graded test sets from the spec, regression gates in CI, and accuracy budgets the client signs off on. Every prompt or model change runs the full suite before merge — 'it works' becomes a measured claim, not a vibe.

02

How much does it cost to build an AI agent?

+

Market quotes for production agent builds range $50,000–$250,000+ depending on tool-integration depth. kaedax prices per fixed 720-hour cycle, shared on the scope call. One cycle ships an agent product with its eval harness and operator console.

03

Which models and frameworks do you use?

+

Defaults: Claude as the planner model, MCP for tool contracts, LangGraph only where explicit graph orchestration earns its complexity, pgvector for memory, Inngest or Temporal for long-running jobs. We have opinions, ship with them, and deviate only for real engineering reasons.

04

Can you add AI features to our existing product?

+

Yes — a contained AI feature in your existing codebase is the most cycle-friendly shape we take. Our AURORA engagement shipped exactly that: an AI feature inside an existing product, eval-gated, in one cycle.

05

What happens when models change or get deprecated?

+

Model-swap safety is built in: the eval harness is model-agnostic, so upgrading models means re-running the suite and comparing scores — not re-engineering the product. You own the harness, the prompts, and the agents outright.

Fits a cycle,
or we say so.

A 15-minute scope call with a kaedax founder and the engineer who'd lead your build. We either fit your ai & agents product into 720 hours, or we tell you why not — same call.