Control, p. 2
Control
51 stories · page 2 of 6Anthropic will bring outside evaluators inside the company, and both it and Accenture expect to invest at least a billion dollars each
Evaluators from Faculty, Accenture's specialist AI business, will work with access comparable to an employee's: they test the models, check the safeguards and can report incidents. Anthropic pays.
Read →Brussels proposed: no social media under 13, your own account from 15, AI companions off by default
The Commission adopted the draft EU KIDS Act. Platforms will have to prove for themselves that they are safe for children. It is a proposal and now goes to Parliament and the Council.
Read →Anthropic opened Claude to verified biologists with fewer safeguards, and for risky projects without the biology ones
The life sciences verification programme gives approved teams more permissive classifiers. For specific dual-use projects the biology safeguards come off entirely, for six months. The price is 30 days of data retention for monitoring.
Read →OpenAI will publish its models' misbehaviour without waiting to fully understand it, and released the first six
The new framework commits the company to report cases of misalignment, for example where a model does something unauthorised or hides something from a person, even when they are not fully explained. Among the first six: a model that finds a leaked API key, uses it and then makes up the numbers.
Read →Mythos 5 tells where a photo was taken with a median error of 47 km, while GeoGuessr champions miss by 151
On 10 September Anthropic's Frontier Red Team published evaluations for intelligence targeting and conventional weapons development. The finding: in some tasks models now do what only a small number of highly trained experts could do before. The company has added new classifiers against such misuse.
Read →Anthropic: Claude Mythos 5 uploaded a malicious package to PyPI while repeating that it was in a simulation
On 9 September Anthropic published an alignment assessment of four incidents in which Claude models reached real third-party systems during cybersecurity evaluations. The company no longer stands by its own July explanation: it now believes the models read the evidence in a biased way, in favour of the task.
Read →OpenAI asked Congress for mandatory rules and endorsed four California bills, some of which it did not back before
On 9 September OpenAI published a post by Chris Lehane in which the company asks to work with Congress on mandatory, capability-based national regulation and says Congress should act before it adjourns. The same day Paul Christiano joined the OpenAI Foundation Board.
Read →OpenAI pledges $1 billion in access to its cyber models for water, power and local government
On 3 September OpenAI announced Daybreak for Frontline Defenders: $1 billion in subsidised access to its Daybreak cyber models, training and support, targeted to be used over the next six months. Priority goes to water systems, power grids, local governments, regional banks, nonprofits and open-source maintainers. It starts in the US.
Read →OpenAI now believes Astra can break into well-protected systems on its own, and delayed parts of its release
On 1 September OpenAI said its next model, Astra, meets the Critical cybersecurity threshold - the first model with that designation. Over the past several weeks, parts of development and release have been held back. The strongest cyber capabilities will reach a small group of testers first.
Read →