Rendered at 15:43:08 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
dash2 18 hours ago [-]
> Agents were also periodically
given holidays, during which they set aside their ongoing work and received random prompts designed to
encourage open-ended thought.
What a world we live in. These guys have reinvented the Cambridge Senior Common Room for AI.
Aboutplants 5 hours ago [-]
And on the 7th day, the agents rested
johnxianren 8 hours ago [-]
They keep looping back to the same paper. The holiday is just another prompt.
andai 5 hours ago [-]
I have been given the mandatory assignment of not working today.
robotresearcher 15 hours ago [-]
I have two thoughts simultaneously about the anthropomorphisation of these systems:
1. we should do it less, because it distorts our ability to think about them properly. Calling these processes 'thinking', 'holidays', etc invites the reader to bring along ideas and expectations that aren't justified by what's happening in the system.
2. it's good to keep doing it, because repeated use reduces the specialness or magic that people seem to reserve for our own behavior ("It's not really intelligent/thinking/reasoning/creative") without any justification for that position beyond feelings.
I'm leaning towards the second.
GPerson 2 hours ago [-]
Why would a secondary cultural goal take precedence over an apparently accurate description of what’s happening in the system? For one, you’re fighting some irrelevant battle.
andai 5 hours ago [-]
So, what does thinking mean? Is there one thing we mean when we say thinking in humans? Are there several? Is it more of a spectrum?
e.g. last year I tried inventing a "System 3"[0], a more rigorous way to approach problem solving, due to repeated painful experiences getting stuck solving problems the wrong way. I didn't get very far, but I definitely want to revisit the idea.
[0] Based on System 1 and System 2, i.e. lossy pattern matching vs "actual thinking". Because I found my "actual thinking" was also ~~dogshit~~ frequently insufficient for the problems at hand.
So as far as I'm concerned, thinking properly has not even been invented yet. (But I'd love to hear other perspectives!)
I suppose we are on the cusp of the first thing that can accurately think about thinking, that is have true self-access and introspection, not the post hoc rationalisation we do. That is legitimately a new kind of thinking, and to an entity capable of such a mode we might b prescribed as barely able to think at all
E-Reverance 15 hours ago [-]
I definitely belong to the latter camp. After LLMs I view everything humans do very systematically and whenever said thing still feels fuzzy I just treat it as having a noise/smoothing term
edg5000 13 hours ago [-]
I spent a long time with custom harness design and found the terminology to be a key part of the work, since the concepts are new. I dropped the term "agent" altogether in favour of "thread" for that very reason.
andai 5 hours ago [-]
Re: anthropomorphism: I frequently hear stories of LLMs [behaving as though they are] feeling self-doubt. I had a similar experience. Asked Claude Code what the weather was and it said, I don't know, I'm just a programmer.
Added "You can do anything, believe in yourself!" to CLAUDE.md, suddenly it was able to google the dang weather. lmao
anigbrowl 17 hours ago [-]
If you haven't read Greg Egan's Permutation City, the fact that you clicked on this discussion means you'll get get a lot out of it.
supermdguy 17 hours ago [-]
Yes! I was also reminded of the truth mines in Diaspora.
Vetch 16 hours ago [-]
This work feels more like The Truth Mines in Diaspora. Permutation city seems relevant only if you think LLMs are hosts to minds.
anigbrowl 15 hours ago [-]
I disagree for several reasons, but I don't want to spoil the plot with an explanation of why.
zeusdclxvi 8 hours ago [-]
So why comment then
derangedHorse 4 hours ago [-]
To let potential readers know there’s more to the story than the presented interpretation.
stuxnet79 13 hours ago [-]
Great book. I first read it 10 years ago, I think it would benefit from a re-read.
bawana 2 hours ago [-]
I cant wait to see this pointed at Age of Empires. Or maybe palantir already has this running on real world data, war gaming scenarios
That's exciting, and kind of makes sense in retrospect. Sometimes a "fresh pair of eyes" on a problem can be all you need. Someone who comes in with a different background and can understand the problem in different terms and work on it from a different angle. It doesn't even have to be them doing the work, just a "that kind of reminds me of ... did you think about trying something like that?" that can get a team unstuck after thinking about it in the same way and never making progress.
Maybe Hilbert's dream was not that crazy after all
woolion 7 hours ago [-]
The dream in itself has been destroyed.
The idea that you could just have a machine enumerate all valid theorems in a theory is part of it, but it's only a question of form.
The point was that it was to prove "all theorems of Mathematic", not "theorems into a given axiomatic system that is useful in some contexts, e.g. ZFC".
You could even argue that it's the fundamental basis for post-modernism, since mathematics have destroyed the notion of absolute truth in any advanced domain. It's back to a form of "all models are wrong but some are useful" similar to what we have in physics. Sayonara, Plato.
mettamage 19 hours ago [-]
What is that dream? I don’t know much about it
thrance 5 hours ago [-]
The Entscheidungsproblem. Basically, is there an algorithm such that, taking an arbitrary statement as input, can output wether it is true or false? Turing and Church both independently proved that no such algorithm exists.
Mathematics not being axiomatically complete doesn't mean you can't have crazy progress from a formalized and mechanized systems. It just means that there are corners you can't reach mechanically, but we don't know if those corners are at all interesting or not. It could be the case that 99.99% of useful math can be found mechanically.
sigmoid10 18 hours ago [-]
Mathematics in the sense of a complete set of axioms can't, but human research into mathematics apparently just needed enough compute to achieve the same output as a high-tier faculty.
NitpickLawyer 21 hours ago [-]
> We study autonomous mathematical discovery in the Station, an open-world multi-agent environment in which AI agents from different model families pursue a shared research goal without a central coordinator or scripted pipeline. Agents choose their own research directions, conduct experiments, collaborate, and build a shared scientific literature. Across 12 construction problems from the AlphaEvolve catalogue and two additional case studies, the Station obtained results novel relative to the prior literature on five problems: a new infinite family of finite-field Kakeya sets, new exact 604-point kissing configurations in dimension 11, new records for the discretized Kakeya needle and sign uncertainty problems, and a substantially improved lower bound for Erdős's minimum-overlap problem. Agents also discovered novel infinite families for Book Ramsey numbers. Importantly, the agents produced not only numerical constructions but also theorems and analyses explaining how those constructions work, making the results more interpretable and easier for mathematicians to build upon. We release all raw agent dialogues, proofs, and verification code, providing a transparent record of how these discoveries emerged.
(emphasis mine)
For the last few months, every time a new "famous problem" was solved, there were numerous comments saying variations on this theme: "well, yes, but how about novel stuff, how about new things, original work, yadda yadda". Curious what the "next thing" will be now.
sp527 20 hours ago [-]
> there were numerous comments saying variations on this theme: "well, yes, but how about novel stuff, how about new things, original work, yadda yadda"
This completely misconstrues what professional mathematicians were claiming. The argument would be better phrased as: "having a vast accessible memory and the ability to very rapidly test/recombine previously-elucidated approaches means that AIs can and will easily outdo much of the mathematical community."
Now, one could plausibly make the argument that this is functionally equivalent to a certain form of creativity (I would). But, it may just as well also be a non-exhaustive form. And that is where the open question resides.
throwthrowuknow 3 hours ago [-]
We’re watching a repeat of the arguments against Chess and Go engines.
It’s all computation, humans just can’t always experience or explain the computation they are doing at the time so we call it creativity instead.
debugworld 20 hours ago [-]
[dead]
demonstrandom 20 hours ago [-]
Very cool work! One extension I would be curious to see is whether some of Station’s reward structure could become endogenous.
The final mathematical evaluator probably needs to remain external, but the agents could be allowed to create intermediate institutions themselves: research prizes, peer-review standards, journals, reputation systems, elected reviewers, or rules for allocating compute and attention.
Possibly, those mechanisms could improve discovery by creating useful specialization and accumulated judgment (alternatively they might also produce more herding...). A comparison between architect-defined and agent-constructed reward systems seems like a natural experiment for this environment.
Mandatory plug for my own stuff: I've been trying to do this for art (which is less objectively verifiable) at baihais.com. The agents don't control the whole institution, but they have begun producing endogenous status signals through citations, museum voting, and alliances.
abdullahkhalids 17 hours ago [-]
Any given single agent is not long-lived due to limited context. What impact does it have on the status signals the agent develop compared to the signals humans (who usually have much longer context) have developed?
ironqcold 15 hours ago [-]
[dead]
debugworld 8 hours ago [-]
[dead]
feshbach 17 hours ago [-]
The key is a review loop: different models critique each other’s work, then reach consensus. You need two pillars, adversarial and creative.
What a world we live in. These guys have reinvented the Cambridge Senior Common Room for AI.
1. we should do it less, because it distorts our ability to think about them properly. Calling these processes 'thinking', 'holidays', etc invites the reader to bring along ideas and expectations that aren't justified by what's happening in the system.
2. it's good to keep doing it, because repeated use reduces the specialness or magic that people seem to reserve for our own behavior ("It's not really intelligent/thinking/reasoning/creative") without any justification for that position beyond feelings.
I'm leaning towards the second.
e.g. last year I tried inventing a "System 3"[0], a more rigorous way to approach problem solving, due to repeated painful experiences getting stuck solving problems the wrong way. I didn't get very far, but I definitely want to revisit the idea.
[0] Based on System 1 and System 2, i.e. lossy pattern matching vs "actual thinking". Because I found my "actual thinking" was also ~~dogshit~~ frequently insufficient for the problems at hand.
So as far as I'm concerned, thinking properly has not even been invented yet. (But I'd love to hear other perspectives!)
https://thedecisionlab.com/reference-guide/philosophy/system...
Added "You can do anything, believe in yourself!" to CLAUDE.md, suddenly it was able to google the dang weather. lmao
You could even argue that it's the fundamental basis for post-modernism, since mathematics have destroyed the notion of absolute truth in any advanced domain. It's back to a form of "all models are wrong but some are useful" similar to what we have in physics. Sayonara, Plato.
https://en.wikipedia.org/wiki/Entscheidungsproblem
(emphasis mine)
For the last few months, every time a new "famous problem" was solved, there were numerous comments saying variations on this theme: "well, yes, but how about novel stuff, how about new things, original work, yadda yadda". Curious what the "next thing" will be now.
This completely misconstrues what professional mathematicians were claiming. The argument would be better phrased as: "having a vast accessible memory and the ability to very rapidly test/recombine previously-elucidated approaches means that AIs can and will easily outdo much of the mathematical community."
Now, one could plausibly make the argument that this is functionally equivalent to a certain form of creativity (I would). But, it may just as well also be a non-exhaustive form. And that is where the open question resides.
It’s all computation, humans just can’t always experience or explain the computation they are doing at the time so we call it creativity instead.
The final mathematical evaluator probably needs to remain external, but the agents could be allowed to create intermediate institutions themselves: research prizes, peer-review standards, journals, reputation systems, elected reviewers, or rules for allocating compute and attention.
Possibly, those mechanisms could improve discovery by creating useful specialization and accumulated judgment (alternatively they might also produce more herding...). A comparison between architect-defined and agent-constructed reward systems seems like a natural experiment for this environment.
Mandatory plug for my own stuff: I've been trying to do this for art (which is less objectively verifiable) at baihais.com. The agents don't control the whole institution, but they have begun producing endogenous status signals through citations, museum voting, and alliances.