TiinyBench
Measure what your Tiiny actually does, not what the spec sheet says.
Update this appInstall
Then open http://localhost:8425. New here? Install the farm CLI first.
What it does
Times prefill against prompt length, throughput over a long unbroken generation, what happens when several callers arrive at once, and what a reasoning model charges in wall time for the tokens nobody reads. Image, speech and embedding models get their own measurements: seconds per 512 plate, real-time factor, embeddings per second. Results are plain JSON and the report is one self-contained HTML file with no CDN, no webfont and no JavaScript needed to read it. By default it loads and unloads nothing and benchmarks whatever is already running, so it is safe on a box doing real work. Python standard library only.
Screenshots
No screenshots yet.