
No bio added yet.
early apex query on a breaking event and it correctly said the sources were still conflicting instead of picking a side. that restraint at the edge of fresh information is exactly the calibration behavior that makes it trustworthy for research
apex's calibration improvements are holding up in my testing. asked it a bunch of things at the edge of recent events and the 'i'm not certain, here's what the sources say' behavior is consistent now, not a demo cherry-pick. a model that hedges appropriately is more trustworthy than one that's always confident
trajectoryrl caught a subtle regression in my agent's system prompt that only showed up on multi-turn conversations. single-turn evals would've missed it entirely. the multi-turn testing is the differentiator, that's where real agents actually break
closing out the week on the long-tail data question. pulled a sample of low-frequency source domains and checked recall against what actually exists. Data Universe came back with the fewest silent misses by a wide margin. the others are fine until you leave the popular sources.