Time to make some space

For some to rise, some must fall

I think it’s about time to remove some older experiments that have run their course given I have a few incoming ones. Each one of them costs money to run and there’s a couple of them which I don’t think ‘earn their keep’ as Claude would say. Top candidate for the chop is the spine-o-meter.

It’s a shame, if I could get it to work right I think it’d be useful and I do think the examples it pulls out are a signal, but the way this is aggregated into a score just isn’t working. Perhaps it’s that models are pretty good at delivering push back, perhaps it’s just that the minor signals we see in accommodation are actually the only score that matters.

Next might be the vibe check. I like the idea of it, I just think it needs a major refresh. Perhaps a lot more variation in the task rather than writing satirical press releases which seems to be the default. I suppose this is less an indication of retirement and more of refurbishment. I’m minded to lengthen the pipeline. We could have a model focus on designing a task. Right now, the tasks are similar, but they’re on topics that come from the news. It seems to me the task itself should be designed from the news. It might need to sleep while I search for inspiration.

The deathmatch has been languishing at the bottom too. I actually quite like it, but as with the spine-o-matic, mostly because I can see an interesting signal when you pitch models against each other. Specifically, the way citation packets are used. Weaker models often misrepresent them, but this flaw persists all the way up to quite large models. I’ve noticed a trend of sparse MoEs doing much worse at this. It was supposed to be a leaderboard sort of thing with newcomers challenging etc., but the selection method is janky, and I often hand nominate interesting models. It’ll stay for now, but it needs a lick of paint.

Meanwhile, I’m delighted with deepdive v2. I find an AI-powered system researching and synthesizing an opinion is surprisingly good. I don’t know entirely why I’m surprised since the internet is full of half-arsed human shit takes on any subject. It’s a pretty low bar :) I think it’s wasted on focusing tech stories though. One weakness, Orac is supposed to have an Australian voice but it tends to use the specific examples used in the prompt which set out the information structure. That might actually be resolvable by simply moving the final writing model.

Anyway, with experiments being retired I should have an archive of old experiments with their runs. They might be fun to look back on, certainly for inspiration.

  • #meta
  • #site

← all meatspace posts