A model never reaches a user alone. Defaults, labels, context, review steps, and the opportunity to correct an answer all determine what its underlying capability becomes in practice.
This is why two products built around similar models can produce notably different outcomes. One treats uncertainty as information to expose; the other conceals it behind a polished response.
Evaluation should include the whole encounter. The interface is not decoration around the system. It is part of the system being measured.