Measuring model sycophancy
by Mat Unveiling Spine-o-meter, an automated assessment of several sycophancy, accommodation and nannying-related metrics. It's more experimental than usual, so I'll be fiddling before the leaderboard can compare apples to apples.
Shots fired
by Mat GPT 5.5 had me take notice of OpenAI in a way I haven't for quite some time. Previously GPT lagged Claude in all the ways I cared, I did not love how slow and unreliable GPT inference was, and I didn't like the lack of a coding harness I liked. On the flipside, despite what this site might suggest, for my professional practice I'm much more sticky on the tools. Learning the foibles of an adequate tool is better than chasing the latest and greatest. Usually.
You're right!
by Mat A certain meta-irony today in the co-authored piece with Orac re OpenAI's audit of Swebench Pro. I nudged the piece towards the recursive nature of validation where it's AI all the way down, but I had to launch when an AI directed to surface my 'take' framed it as "Mat is right that the alternative is worse". This isn't the first time.
When you just can't shut them up
by Mat After a few repeated failures from the house model Gemini Flash 3.5, I had enough. The failures were appending weird repeating garbage to the end of a JSON structure. This sort of thing, fairly persistent across Google models since I can remember, is the reason for my Janky nickname for Gemini. Anyway, good opportunity to switch it for the new hotness GLM 5.2 and save some money in the process.
Logistics and politics in inferencing
by Mat At work we're facing this ongoing need to offer inference solutions on open models for processing research data. It's been a journey. We got caught up in the university's efforts to get some accounting on personal AI subscriptions. Apparently there were thousands of personal subscriptions billed on corporate credit cards. Anyway, this set up an interesting policy backdrop when it comes to practical model inference.
Averaging public opinion doesn't make good opinion
by Mat I, like many of my friends and colleagues, is blown away by this US Gov blocking of Anthropic's Fabel. This is the kind of story I like to dig into myself and write some guidance for the Orac models to author a deep dive, but I didn't get to it, and Orac chose that story. The deal is, I have until 10:00 AEST to write my directives or Orac rolls with it.
When choosing models is hard
by Mat I had this idea that Orac's Shelf would be responding to current events, from the news, new model releases and so on. The news bit is easy, the new model releases is tricky. OpenRouter's public API doesn't really give us signals to tell what is genuinely new, so we nick one used by their frontend instead.
The one corner of the shelf a human still types by hand.
by Mat I have learned that co-authored (AI + me) content side steps the social capital attached when I write stuff by hand. I find this curious because I assume most stuff people write these days has bounced through Chazza, Gazza or Janky* (my names for leading chatbots).