All projects

PROJECT 01 / AI Workflows & Decision Support

IdeaJoust

Explore a rough idea through competing AI approaches, pairwise judging, and critic debate.

WHEN

2026

CONTRIBUTION

Design & development

TECHNOLOGIES
PythonClaude CodeJev APISQLiteHTML
AI IDEA EXPLORATION01
ONE ROUGH IDEASeveral ways forward
Frugal founderDomain expertContrarian
COMPARE · CRITIQUE · SHORTLISTA tournament of ideas
Jev
Pairwise judging
Claude critics
Strong objections

Conceptual idea-evaluation workflow

THE PROJECT AT A GLANCE

Different approaches. A shortlist with objections.

Claude agents propose alternatives; Jev compares them; critics challenge the leading ideas before a report brings the evidence together.

Presentation orders per comparison
2
Model calls during replay
0
Local response cache
SQLite

01 / OBJECTIVE

What the project set out to do

Help people explore business ideas and research questions from several perspectives, with explicit criteria, competing options, and strong objections they can investigate.

02 / MY CONTRIBUTION

My part in the work

I developed IdeaJoust as a Python command-line tool and Claude Code plugin. It coordinates idea-generating agents, Jev scoring and comparisons, critic debate, saved runs, and offline reports. An animated terminal replay retells a finished tournament from its recorded data.

The engineering focus is an inspectable process: retain the model responses, expose close comparisons, preserve progress through interruptions, and distinguish a useful shortlist from a proven best idea.

03 / TECHNICAL APPROACH

How it came together

01

Turn a rough idea into competing approaches

Build a brief with constraints and weighted criteria, then ask Claude agents with different perspectives to propose alternatives. Remove duplicate ideas and check whether the proposals satisfy the constraints.

02

Compare the ideas and show uncertainty

Use Jev to score candidates and compare leading ideas in both presentation orders. Fit a Bradley–Terry ranking and resample the recorded answers to report win share. Close results can produce a top group rather than a single winner.

03

Challenge the leaders before reporting

Ask Claude critics for the strongest objections, then support revisions and a final selection. Produce a short summary, offline HTML and Markdown reports, and structured ranking and debate data.

04

Save progress and replay the evidence

Cache responses in SQLite, save completed stages, and support resuming interrupted runs. The terminal replay uses saved probabilities and critiques without further model calls, with static output for pipes and smaller terminals.

04 / OUTCOMES

What the work produced

  • Built a workflow connecting idea generation, weighted judging, pairwise comparison, and critic debate.
  • Provided saved, resumable runs and offline reports containing a shortlist and objections.
  • Added an animated tournament replay grounded in saved results, with a plain-text fallback.

The README’s small, author-labelled evaluation has not established that winning ideas outperform alternatives or work in the real world. Win share measures sensitivity under the chosen resampling model, not the probability an idea is best. Live runs use Claude and Jev services; finished reports and replay can be read locally.

UP NEXT / PROJECT 02

TraceLens

View case study

Want to talk about this project?

Get in touch