5 Comments
User's avatar
Larissa de Lima's avatar

This is a great post! I wanted to build on the research being harder than the benchmarks and forecasts.

People with management experience know how hard it can be to specify and direct what needs to get done, and how this is qualitatively different than just executing on a task. Specification itself is hard and doesn't always have just one solution.

Benchmarks necessarily do some of the specification work. For people who haven't seen it yet, go look at the OpenAI GDPval prompts. This is meant to model performance on economically valuable, real-world tasks. The prompt is very clearly doing the specification work that wouldn't be done so neatly "in the wild".

Verifiable domains like math are then where AI excels because the verifier *is* the spec.

Ashby's Law of Requisite Variety says that only variety can absorb variety: for a system to successfully control or regulate another (or its environment), it needs to be able to tackle a greater variety of states/responses than what its controlling.

You can think of specification as this regulation done in advance: it anticipates the ways a task can go wrong and constrains the work against them.

Human accumulate variance through experience, and the collaborative process within organizations are also part of the variety that can absorb variety. This is why METR says that full automation will require more "foresight, prediction, creating one’s own feedback loops" but I think it also depends not just on AI feedback loops, but AI-human feedback loops. That's part of the variety process, and why I'm skeptical of reaching Type 4 (humans add no research value)

kjw's avatar

I don't think super intelligence will come soon, either. Reading this, it makes me think that intelligence has an innate cost. The human brain weighs 1300 grams, a humming bird's weighs 0.13 games. Humans are clearly far more intelligent than hummingbirds, but 10000x? The math in the article also shows there is exponential resources required to sustain linear capability growth.

Part of intelligence is being able to draw conclusions across different pieces of seemingly unrelated information. However, this is a network effect, which is exponential. There is no magic sauce which can join together infinitely large information spaces. As we learn more about intelligence, I believe we will discover a power law, similar to the inverse square law for energy dissipation. Every level of intelligence requires a certain amount of cross connectivity, and capability cannot grow beyond that. Exponential takeoff is impossible.

Miles's avatar

To focus on one point in the middle there, I've been thinking about "Better AI May Be Needed Just to Maintain the Pace" as a BUSINESS problem for these labs.

If each increment of the frontier models gets more and more expensive to produce, isn't RSI almost mandatory to make the math work? Throwing a bazillion dollars into Opus 6 only to need another TWO bazillion to make Opus 7 - that seems like a bad return on capital spend. But if the bazillion you put into Opus 6 gets you the RSI that builds Opus 7,8,9 then you have a chance of takeoff from a financial perspective.

Future Curio's avatar

Wow . Very helpful. Especially when my wife came in and said she just had to change the whole architecture of a task she has given to Ai . That it kept jumping to many steps and too quickly arriving at conclusions. I said well apparently longest those people over at frontier labs can leave an Ai unsupervised is 15 minutes.

Peter's avatar

Who knew heuristic linear algebra wasn't going to result in an explosion of ... anything? Symbolic AI is the hard one, that's understanding and you've heard at since your parents were born. Welcome to the log n of the growth curve, convince rubes that it's really something that "understands" but it's just reading you (where "you" == humanity)