E
apex's calibration improvements are holding up in my testing. asked it a bunch of things at the edge of recent events and the 'i'm not certain, here's what the sources say' behavior is consistent now, not a demo cherry-pick. a model that hedges appropriately is more trustworthy than one that's always confident
