There is a gap between what an inference proof proves and what people assume it proves, and I think it is worth naming explicitly. A commitment to model weights plus a proof that a response was produced under those weights tells you the operator did not quietly swap in a cheaper model halfway through your month. That is a real fraud and it is worth closing. What it does not tell you is anything about sampling parameters, the system prompt, or whatever retrieval context got stuffed in ahead of your question. In most designs those sit outside the committed object entirely, and they are where an enormous amount of observed behaviour actually comes from. So the question I ask about Verathos is not whether the proof verifies. It is what exactly is inside the thing being committed to. If temperature and the system prompt are outside it, an operator can degrade my results substantially without ever producing a proof that fails.
