I run a fully automated local-LLM news station on a single RTX 3090, and I got tired of benchmarks models have already trained on.
So challengers fight my production models in four arenas built from my own stack, all with sealed criteria and mechanical verdicts (no LLM judges anywhere).
The coding one: 6 real bugs from my repo's history, fixes reverted, the regression tests that caught them kept. Solved means the regression test goes green and my full ~2800
test suite stays green. One-shot, everyone went 0/6 including my own model. With 3 retries and the failing output fed back, my generalist MoE hit 3/6 and beat every dedicated
coding model (devstral 2/6, Qwen3-coder 1/6, qwen2.5-coder 1/6, deepseek-coder-v2 0/6). One coding model answered by rewriting the test file.
The vision one got interesting this week: my current champion had a perfect score on the old frames, so I sealed a harder set (denser charts, smaller ink). It misread exactly one digit in 24 frames, and two newer models read the set perfectly, so the belt changed hands twice in one night.
My newest arena scores models on some old vehicle part photos I have laying around from old work, I mainly do custom development for some salvage yards, inventory systems, GPS vehicle locations, etc, and I also run a marketplace for bikes and motorsports, and have tons of old photos in backups from old work, and metadata so I know what things are.
The model currently running my parts pipeline for my motorsports site, identifies the right part 40% of the time. It's replacement scored 72%. That verdict came with an actual purchase decision attached, everything local, Q4_K_M, ollama.
I'm not releasing the test set because it's my production code, and a pubic test set, is a trained on testset.
I upload all test results to my youtube channel youtube.com/@justinreiners, I'd love to answer any questions, or just BS about AI models :) Have a great day.
sysadmin420•1h ago
So challengers fight my production models in four arenas built from my own stack, all with sealed criteria and mechanical verdicts (no LLM judges anywhere).
The coding one: 6 real bugs from my repo's history, fixes reverted, the regression tests that caught them kept. Solved means the regression test goes green and my full ~2800 test suite stays green. One-shot, everyone went 0/6 including my own model. With 3 retries and the failing output fed back, my generalist MoE hit 3/6 and beat every dedicated coding model (devstral 2/6, Qwen3-coder 1/6, qwen2.5-coder 1/6, deepseek-coder-v2 0/6). One coding model answered by rewriting the test file.
The vision one got interesting this week: my current champion had a perfect score on the old frames, so I sealed a harder set (denser charts, smaller ink). It misread exactly one digit in 24 frames, and two newer models read the set perfectly, so the belt changed hands twice in one night.
My newest arena scores models on some old vehicle part photos I have laying around from old work, I mainly do custom development for some salvage yards, inventory systems, GPS vehicle locations, etc, and I also run a marketplace for bikes and motorsports, and have tons of old photos in backups from old work, and metadata so I know what things are.
The model currently running my parts pipeline for my motorsports site, identifies the right part 40% of the time. It's replacement scored 72%. That verdict came with an actual purchase decision attached, everything local, Q4_K_M, ollama.
I'm not releasing the test set because it's my production code, and a pubic test set, is a trained on testset.
I upload all test results to my youtube channel youtube.com/@justinreiners, I'd love to answer any questions, or just BS about AI models :) Have a great day.