AI self-validation
You're right!
A certain meta-irony today in the co-authored piece with Orac re OpenAI’s audit of Swebench Pro. I nudged the piece towards the recursive nature of validation where it’s AI all the way down, but I had to launch when an AI directed to surface my ‘take’ framed it as “Mat is right that the alternative is worse”. This isn’t the first time.
On one hand, mostly my ‘takes’ aren’t anything particularly controversial, but it’s it’s pretty weird that the model frames agreement with something it was more-or-less directed to conclude. Weird and cringeworthy actually, particularly since elsewhere - like a lot of people I’m sure - I need to insert custom instructions ensuring AI pushback against my ‘astute’ ideas. I wanted to get a T-shirt made saying “Claude says I’m astute”.
Okay, so the boring fix is a directive, don’t express agreement or disagreement with me, consider the evidence, present the opposing view if it is attested by your research (the deep dive is basically an agentic loop researcher) etc.
But you know, as with elsewhere and those custom instructions, I find the greater value in the pushback, or even a decent strongmanning of the best pushback. Claude Opus does this exceptionally well which is why it is my partner in crime for anything important.
The house model is GLM 5.2 as I write this. I can’t help but wonder about Chinese cultural values and AI. China, like many other countries, has a higher acceptance of the difference in power between people, I’m referring here to Hofestede’s Power-Distance index. This manifests in all sorts of ways in theory, and in my experience living in Taiwan for my PhD research. The (relative) unquestionability of power, the respect and deference to elders and teachers and so on (my friends in rural Taiwan could not talk about me as a student, albeit a PhD student in my 40s, they would only refer to me as laoshi or teacher). I cast no judgement here, I just note it’s rather a different world view than Australians who are something of an outlier (Gemini tells me 36 out of a world average of 60). This means I pretty much expect to be called a boofhead at all times, and to get away with calling my boss a boofhead too. Maybe not a judge, mind, I mean it’s 36 not zero.
My rambling point here lands on the question: Am I so bothered by something agreement with me, the lack of pushback and hence sensitivity to sycophancy because of my cultural background. I think I’m in the pre-brainstorm phase of wondering if we couldn’t have some kind of sycophancy experiment here on the ‘shelf’, what do you think? Haha, just kidding I know it’s just me and bots reading this.