Filipe Pawlik Leite
I make AI agents stop lying.
I build the harness around AI agents — the gates, evaluation and orchestration that decide whether an agent's output is good enough to ship. Then I sit with the customer, find the real problem, and stay until the thing works.
A gate that
refuses.
Anything an agent writes goes through a gate before shipping: typecheck, a real test suite and — on web — a build that produces a servable artifact. Nothing ships otherwise. Flip the switch and break the code. The gate below is the same logic, running here in your browser.
Press run. The gate does not know which variant you picked.
Proven honest with shadow tests: known-bad code is refused in front of whoever opens this page. When the gate fails, the error goes back to the agent, which retries with backoff — up to a hard cap.
Every number here
was counted.
If a figure on this page has no source, it should not be on this page. Small samples travel with their sample size.
Routing by hit rate,
not by vibe.
Measured accuracy decides who gets the next task, sample size attached. Pick a job and watch which agent the router would call, and why — the same rule it runs in production.
A hit rate is only shown where n is large enough. The rest appear with their n — too few runs to route on.
Hospitals, blood banks,
governments, Chevrolet.
Systems that run
without me.
For every layer, the same three questions: what it really costs, how good it actually is, and whether it holds when the volume grows. Then I measure — and the decision is what survived measurement, not the favorite candidate.
Let's find your
real problem.
filipe@netcks.com
Bombinhas, Santa Catarina, Brazil. UTC−3 — full overlap with US business hours. English, Spanish and Portuguese, all fluent, all working languages.