Of course that depends on having controls that can't be circumvented which is a big if
The whole point of intelligence is that it's generalizable. If you constrain it to some controlled things, then it ceases to be useful. its incompatible. the whole incentive with ai is to let it do whtv it wants.
Actual title: "New MIT Algorithm Meets Every Hard Constraint in Simulated Tests"
My (maybe naive) question is if we only check the final result then isn't it already too late and possibly the safety rules have already been irreversibly violated? It gives the example of a robot arm avoiding obstacles while still finding the shortest path, but if we only check the correctness at the end, then isn't it possible that it already collided with an obstacle?
You would put the check before it actually does the thing. It's at the "end" of the process of figuring out what it wants to do, not the end of fulfilling the request or prompt
HardFlow seems to be a strategy for nudging the model in the right direction while giving it more freedom. Only applying the constraints at the end is a mischaracterization from what I can tell.
Hope this approach gets well tested and sees good results so we have a shot at human governance.
> Our key insight is to leverage numerical optimal control to steer the sampling trajectory so that constraints are satisfied precisely at the terminal time.
Doesn't seem so "fool proof" to me as where the inevitable media spin will take it. Then, how do you know "where" to steer weights? "Safe" has no agreed upon definition
1. My agents do not take direct action, they run programs.
2. Programs are not LLM hits/real-time output; they don't MAX tokens the MAX determinism.
3. Programs are logistical wrappers for Protocols where the guardrails are (RLVR.ai)
4. Policies for generating programs are democratically governed re: fec.dev - they get voted on
5. Elected-HITL implements the policy pipelines and ontological abstract intents [and their maps]
I would be really interested to learn about the Gov. models that others are using but this seems to be something that linked0-in (which i loathe) discusses (in the most pedestrian/luddite) terms more than HN.
What are the rules? Or does sharing the rules present security problems, so they're not shared?
What are the ethics axioms?
And why should I trust your ethical framework?
Basically, anything that is leaving the LLM is checked, rather than the internal LLM reasoning process.
ck2•43m ago
hmokiguess•42m ago
dessimus•24m ago
guywithahat•38m ago