Discussion about this post

User's avatar
John C's avatar
3dEdited

Great post, long-time Naam fan here.

I'm an older working R1 research scientist in an AI-adjacent field.

This post rings true to me. And it boils down to one thing IMO, learned through a long career in science. Meaningful intellectual progress, making new insights and understanding, is exponentially hard. That means progress is logarithmic in time-integrated effort.

Popular history gets this wrong... it's all Eureka! a lone genius has a brainwave and the world is changed forever. That is the numerator, the insight. The denominator is how many ideas were abandoned, how many researchers didn't have the insight, how many students helped push, but then flew away to industry, how many research grants ended up making only mediocre papers?

And the language around ASI assumes, all we have to do is build a robot Einstein, pump in some electrons and its Eureka's every day, or a rapidly rising hyperbolic crescendo of Eureka's to a finite time singularity.

To me, that expectation is risable. We have forgotten Gödel.

Every past breakthrough, no matter how large, has just pushed back the frontier of human knowledge a little more.

The frontier labs are not building AI Einstein's that can make breakthroughs. We are building tools for an AI-powered scientific apparatus, with humans embedded in it for at least a while longer. And that scientific apparatus will make the same plodding, occasional small jump forward progress as before. And when AI really helps move the enterprise forward, it will just be a subtle but real trend break on the scientific progress curve from 2020 to 2040.

Larissa de Lima's avatar

This is a great post! I wanted to build on the research being harder than the benchmarks and forecasts.

People with management experience know how hard it can be to specify and direct what needs to get done, and how this is qualitatively different than just executing on a task. Specification itself is hard and doesn't always have just one solution.

Benchmarks necessarily do some of the specification work. For people who haven't seen it yet, go look at the OpenAI GDPval prompts. This is meant to model performance on economically valuable, real-world tasks. The prompt is very clearly doing the specification work that wouldn't be done so neatly "in the wild".

Verifiable domains like math are then where AI excels because the verifier *is* the spec.

Ashby's Law of Requisite Variety says that only variety can absorb variety: for a system to successfully control or regulate another (or its environment), it needs to be able to tackle a greater variety of states/responses than what its controlling.

You can think of specification as this regulation done in advance: it anticipates the ways a task can go wrong and constrains the work against them.

Human accumulate variance through experience, and the collaborative process within organizations are also part of the variety that can absorb variety. This is why METR says that full automation will require more "foresight, prediction, creating one’s own feedback loops" but I think it also depends not just on AI feedback loops, but AI-human feedback loops. That's part of the variety process, and why I'm skeptical of reaching Type 4 (humans add no research value)

31 more comments...

No posts

Ready for more?