Vendors publish benchmark scores the way car makers publish 0-60 times: a useful, standardised comparison point, and also a number that can be optimised for in ways that don’t necessarily reflect how the thing performs on your actual road. A model can score well on a public benchmark and still perform poorly on your specific documents, your specific edge cases, or your specific tone requirements, because the benchmark was never testing for those things.
Public benchmarks are a reasonable first filter for narrowing which models are worth evaluating further. They are not a substitute for testing against your own task with your own data, which is the only benchmark that actually predicts how the system will behave once it’s yours to run.
The organisations that get burned are usually the ones that stopped at the public number. The ones that get it right build a small internal benchmark, a set of real, representative tasks with known correct answers, and re-run it every time they consider a model change, not just once at selection time.