@kromem

kromem@lemmy.world · 5 days ago

Actually, they are hiding the full CoT sequence outside of the demos.

What you are seeing there is a summary, but because the actual process is hidden it’s not possible to see what actually transpired.

People are very not happy about this aspect of the situation.

It also means that model context (which in research has been shown to be much more influential than previously thought) is now in part hidden with exclusive access and control by OAI.

There’s a lot of things to be focused on in that image, and “hur dur the stochastic model can’t count letters in this cherry picked example” is the least among them.

kromem@lemmy.world · 6 days ago

I was thinking the same thing!!

It’s like at this point Trump is watching the show to take notes and stage direction.

kromem@lemmy.world · edit-2 6 days ago

Yep:

https://openai.com/index/learning-to-reason-with-llms/

First interactive section. Make sure to click “show chain of thought.”

The cipher one is particularly interesting, as it’s intentionally difficult for the model.

The tokenizer is famously bad at two letter counts, which is why previous models can’t count the number of rs in strawberry.

So the cipher depends on two letter pairs, and you can see how it screws up the tokenization around the xx at the end of the last word, and gradually corrects course.

Will help clarify how it’s going about solving something like the example I posted earlier behind the scenes.

kromem@lemmy.world · 6 days ago

You should really look at the full CoT traces on the demos.

I think you think you know more than you actually know.

kromem@lemmy.world · edit-2 7 days ago

I’d recommend everyone saying “it can’t understand anything and can’t think” to look at this example:

https://x.com/flowersslop/status/1834349905692824017

Try to solve it after seeing only the first image before you open the second and see o1’s response.

Let me know if you got it before seeing the actual answer.

kromem@lemmy.world · edit-2 9 days ago

The pause was long enough she was able to say all the things in it mentally.

kromem@lemmy.world · 9 days ago

They got off to a great start with the PS5, but as their lead grew over their only real direct competitor, they became a good example of the problems with monopolies all over again.

This is straight up back to PS3 launch all over again, as if they learned nothing.

Right on the tail end of a horribly mismanaged PSVR 2 launch.

We still barely have any current gen only games, and a $700 price point is insane for such a small library to actually make use of it.

kromem@lemmy.world · edit-2 15 days ago

Meanwhile, here’s an excerpt of a response from Claude Opus on me tasking it to evaluate intertextuality between the Gospel of Matthew and Thomas from the perspective of entropy reduction with redactional efforts due to human difficulty at randomness (this doesn’t exist in scholarship outside of a single Reddit comment I made years ago in /r/AcademicBiblical lacking specific details) on page 300 of a chat about completely different topics:

Yeah, sure, humans would be so much better at this level of analysis within around 30 seconds. (It’s also worth noting that Claude 3 Opus doesn’t have the full context of the Gospel of Thomas accessible to it, so it needs to try to reason through entropic differences primarily based on records relating to intertextual overlaps that have been widely discussed in consensus literature and are thus accessible).

kromem@lemmy.world · 15 days ago

This is pretty much every study right now as things accelerate. Even just six months can be a dramatic difference in capabilities.

For example, Meta’s 3-405B has one of the leading situational awarenesses of current models, but isn’t present at all to the same degree in 2-70B or even 3-70B.

kromem@lemmy.world · 5 months ago

DeSantis conveniently didn’t mention Mosque leaders as teaching the kids. I wonder why that was. Does Florida not recognize Islam as a religion? Or did he just not want to point out that was a possibility…