A contest, run and settled
Eighteen model × prompt configurations graded against your golden questions and ranked by cost per correct answer.
ClickHouse AgentArena
In 90 minutes you run your own contest — six models across three prompt strategies, answering real business questions in SQL over ClickHouse — and pick what to ship on evidence: cost per correct answer, not a public leaderboard.
Not a public leaderboard. Your own contest.
Picking a model is the first consequential decision in an agent application, and the one most often guessed. Guess high and you pay frontier prices for capability the task never needed. Guess low and you ship something that quietly gets your real workload wrong.
So you run the contest yourself. A grid of six models across three prompt strategies answers a ground-truthed set of business questions against a live ClickHouse dataset.
Every answer is graded on execution accuracy — did the query return the right result, not the right-looking SQL — and every configuration is ranked by cost per correct answer. That is quality per dollar for your task, which is the only number that can actually settle the argument.
Langfuse is wired in from module 00, not bolted on at the end: it runs the contest as experiments, grades it with a code evaluator and an LLM judge, stores every trace, and follows the winner into production.
What you walk away with
On your own ClickHouse Cloud trial, your own Langfuse project, and one OpenRouter key.
Eighteen model × prompt configurations graded against your golden questions and ranked by cost per correct answer.
A correctness evaluator and an LLM judge running server-side in your own Langfuse project, scoring every run.
Answers compared to cached golden result sets — column position, float rounding, row-set normalisation — so a lucky-looking query cannot pass.
Reviewed answers become new golden questions and the grid re-runs — the whole loop visible end to end in Langfuse.
The winning config live behind POST /ask, traced per session, with 👍/👎 written back as Langfuse scores.
Harness, agent, read-only SQL guard, serving API and arena UI — yours to point at your own data.
The route · ~90 min hands-on
One flow, start to finish: pick a model on evidence, understand it, improve it, then watch it in production.
Connect OpenRouter, ClickHouse Cloud, and Langfuse — tracing wired in before any model is chosen — and seed the arena dataset.
Run the model × prompt grid as Langfuse experiments and crown a winner by cost per correct answer.
Go past “it won”: per-tier accuracy, the llm_judge signal, and individual traces from prompt to generated SQL.
Turn reviewed answers into new golden questions through the annotation queue, then re-run the grid against them.
Serve the winning config behind a real API and watch production traces, sessions, cost, and 👍/👎 feedback.
Self-paced by default. Every module names its starting checkpoint and ends with a verification you can check yourself, so nobody is stranded by the pace of the room.
Before you join
Bring a ClickHouse Cloud trial, a Langfuse project, and an OpenRouter key. Leave knowing which model to ship, what it costs per correct answer, and how to prove it again next quarter.