Experiments
Customer-message benchmark
shogo@copilot_shogoJev shipped three days ago and nobody had benchmarked it. So I did. 60 real-shaped customer messages. Same questions, same conditions, zero retries. Jev vs GPT-4o-mini vs Claude Sonnet 4.5. Jev was right 157 times in a row. The system still failed. Speed first, the least
Original post on X