Even an 8yo has better metacognition, it seems. :-)
> The sky is blue because of something called Rayleigh scattering. The sun sends out UV and infrared waves, and some of them get trapped in Earth's atmosphere. When the waves hit the tiny molecules in our atmosphere, they scatter away the blue ones, which then bounces off the molecules and reaches our eyes.
"filtered to the U.S. elementary-school curriculum", suuure
Perhaps the unexpected response comes from its recall ability. It’s not the personality of a child, just the material a child is exposed to.
(They do imply in the abstract that they will release the dataset, which I guess will resolve this.)
LLMs, incidentally, respond in a similar pattern in my experience.
> It's a cat that has been misbehavin'!
I'm assuming the knowledge doesn't end up as separate "layers".
I'm also reminded of how the human mind develops in distinct stages (e.g. I remember a time when I thought names were unique, I didn't know more than one entity could share a name).
> Me: "What's semiotic crystallography? > Response: "I don't know, what is it?"
Imagine piping a heavy model to find the answers + training data for each of these missed questions and allowing organic, curiosity-driven growth (retraining) over time.
A LLM does not learn topic by topic, it learns everything at once and slowly integrates it in to a single knowledge system.
You're asking a lot from extremely fancy auto complete...
Over the last 3 years I've seen projects where I thought, pretty obviously that's a bad idea. But, because LLMs don't say no and can just be pushed to build it anyway, the people building them might never learn that or learn why.
It's nice to be able to have a quick prototype or mvp. But if we never hit friction or something not working out, we never learn or have to come up with a creative solution.
Now, the LLM might seem incredibly intelligent (relatively speaking) and also creative but let's not forget that all is based on its training data. I simply don't believe it can ever be omniscient or that the companies training it are careful enough when doing so.
There's also a second aspect to it, just in terms of RLHF mechanisms. If you've ever experimented with VLA models (i.e. vision input + text task = robotic arm motion output), they tend to need all the training examples of the robotic arm being motionless removed entirely, otherwise the model simply learns that staying still is rewarded and proceeds to never do anything at all. You successfully train the laziest bot in the universe. I wouldn't be surprised if something similar happens to LLMs, if no is a valid answer, why ever do anything?
I have actually gotten "hey i don't think this is a good idea, here's why" as feedback from at least Opus. It WILL still do it if I just demand stupidity (and hell i've been right, which is another topic entirely) but it has given me more confidence this can be a useful tool in the right spots.
That said I probably don't need the top tiers (metrics at least confirm that) and I'm guessing that's specifically because I was working in coding. Most were worded in a "is this a good idea" framing which probably helped, but at least once I said 'lets use this library/method" and it gave a decent argument on why that was basically redundant without prompting.
I still struggle to see the price point panning out.
aetherspawn•45m ago
_diyar•42m ago
aureate•32m ago
This could be seen as an amusingly extreme example of the fact that if you come up with something and state it condidently enough, a surprisingly large number of people will assume you what you're talking about. Presumably, though, you just mistook the unfiltered (trained on the full data) response for the "Little Learner" one.
aetherspawn•10m ago