Jarvis Arena

About

Jarvis Arena records voice assistants trying real tasks, so anyone can see what works today and what does not.

How a run works

  1. A workflow is a short script: what a person says, line by line, and what should be true after each line.
  2. A synthetic voice (Deepgram Aura-2) says each line, and Jarvis hears it through speech recognition (Deepgram Nova-3), so mishearings are real.
  3. Jarvis is a language model (GLM-4.7-Flash by default) with four browser tools: open a page, click, read the page, scroll. It acts on a real Chromium window, then answers out loud.
  4. After each line the checks run: the address, the answer, what is on the page.
  5. The screen and both voices are recorded. The video, the transcript, the checks and the cost of the run are kept.

All of it runs on Cloudflare: a Worker, one Container per run, Workers AI, D1 and R2. Runs are started by hand; the voice in the videos is synthetic and nobody's real conversation is recorded.

Run it yourself

The code is open source (MIT): github.com/eyalev/jarvis-arena (opens github.com). Deploy it to your own Cloudflare account, write workflows, and compare models.

Contact

Made by Eyal (opens eyalev.com). For anything, including corrections and removal requests, write to hello@kapps.dev, or use the feedback form.