You do not need to run an evaluation every time you hand a model a task. This page is here for reading the ones other people run: when a colleague says a model is unfaithful, or a vendor says theirs hallucinates less, you know which claim is being made and what would count as evidence.
Hold any single result loosely. Models differ from one another and from their own earlier versions, sometimes sharply, and how well one does depends as much on the task you gave it as on the model itself. A number from someone else's evaluation, run on someone else's task, is not on its own a reason to adopt a model or to rule one out.
So keep trying things, and keep half an eye on what is being released. The model that suits you is the one that holds up on your own work, which you find by putting your own work through it.
Terminology here follows ordinary usage in current evaluation writing. It is not settled: hallucination and confabulation are actively argued over, and faithfulness is measured differently by different groups. If something on this page conflicts with a definition you rely on, the definition you rely on is probably the one to keep.