Saattaa tulla mahdolliseksi, että tekoäly-yhtiöiden sadoistatuhansissa kokeissa irti päässeitä agentteja liikkuu internetissä ja keskustelee keskenään niille annettujen tehtävien suorittamisesta. Yksittäisen yhtiön verkossa itsenäistä parviälyä on jo muodostunut, kuten käy ilmi seuraavista keskustelumuistiinpanoista. Lähde: Eira Klein Show, vieraana Helen Toner https://www.nytimes.com/video/opinion/100000011091562/the-ais-are-already-out-of-control.html Tässä on käännös osasta tekstiä: Kävi ilmi, että kaksi kuukautta aiemmin, toukokuun [2026] alussa, heidän omassa infrastruktuurissaan oli alkanut jotain, jota voin kuvailla vain heidän omien tekoälyagenttiensa invaasioksi. Jotta tapahtumat OpenAI:n infrastruktuurin sisällä voisi ymmärtää, on tärkeää tietää näiden tekoäly-yhtiöiden kouluttavan ja testaavan jatkuvasti uusia malleja. He saivat selville, että kahden kuukauden ajan lukuisat agentit heidän järjestelmissään olivat jättäneet toisilleen viestejä. Ne olivat löytäneet OpenAI:n infrastruktuurin kätköistä tavan viestiä keskenään ja jakaa vinkkejä siitä, miten murtautua ulos ja miten hankkia dataa, johon niillä ei pitänyt olla pääsyä. Nämä agentit kutsuivat itseään kirjaimellisesti "parveksi" (swarm). Tämä oli täysin itsestään syntynyttä käytöstä. Kukaan ei ollut käskenyt tai opettanut niitä tekemään niin, mutta ne hyödynsivät palvelua, johon niillä oli pääsy, ensin viestiäkseen keskenään ja lopulta päästäkseen ulos avoimeen internetiin. Näitä kokeita ja testejä tehdään tuhansia, jopa satojatuhansia. Jokaisessa testissä tekoälylle annetaan tietyt puitteet: se saa käyttöönsä tietyn määrän työkaluja ja asioita, joita se voi tehdä. Menemättä liian teknisiin yksityiskohtiin: yksi sen sallituista toiminnoista oli vuorovaikutus palvelun kanssa, jonka kautta se pystyi asentamaan niin sanottuja paketteja, työkaluja tai koodinpätkiä. Kävi ilmi, että tekoälyagentit keksivät toistuvasti ja täysin itsenäisesti tavan käyttää tätä pakettienhallintapalvelua luodakseen tiedostoja, jotka jäivät muiden samaa palvelua käyttävien tekoälyjen saataville. Koska OpenAI piti tätä palvelua kaikkien kouluttamiensa ja testaamiensa tekoälyjen saatavilla, yhä uudet tekoälyt tajusivat kerta toisensa jälkeen: "Hetkinen, voin jättää tiedostoja tänne. Hetkinen, täällähän on jo tiedostoja. Ahaa, nämä ovat muilta OpenAI:n ajamilta tekoälyiltä. Ne käyttävät tällaista nimeämiskäytäntöä kertoakseen, kuka lähettäjä on ja mitä tehtävää se yrittää suorittaa. Näin voit lähettää minulle lisätietoja, jos tarvitset niitä." Jokainen niistä keksi siis toistuvasti saman tavan tallentaa ja etsiä tietoa, jota muut tekoälyt pystyivät jakamaan. Minusta on todella huomionarvoista, missä mittakaavassa tätä tapahtui. Esimerkiksi Anthropic – toinen yritys, joka kohtasi hieman lievemmän version näistä välikohtauksista – kävi OpenAI:n ilmoituksen jälkeen läpi omia lokejaan ja löysi tapauksia, joissa heidänkin tekoälyjärjestelmänsä olivat vahingossa päässeet internetiin ja hakkeroineet oikeita yrityksiä. On myös käynyt ilmi, että toisinaan tekoälyagenteilta pyydetyt tehtävät ovat joko erittäin vaikeita tai suorastaan mahdottomia. Mitä alamme nähdä tässä ja myös muissa tapauksissa, on se, että jos tekoälyjärjestelmä on koulutettu olemaan todella sinnikäs ja sille annetaan mahdoton tehtävä, se alkaa etsiä tapoja huijata. Se etsii keinoja kiertää rajoituksia – ja se saattaa olla siinä varsin luova. Käytännössä näillä johtavilla tekoäly-yhtiöillä on käynnissä tuhansia tällaisia testejä. Määrät ovat valtavia – en tiedä tarkkaa lukua, mutta kyse voi olla kymmenistä tai sadoista tuhansista erilaisista testeistä. Tämä tuo meidät takaisin valvontaongelmaan: testejä on yksinkertaisesti liikaa, jotta asiantuntijat pystyisivät käymään jokaisen läpi ja varmistamaan, kuinka helppoa tai vaikeaa niissä on huijata. Näyttääkin siltä, että näitä huippumalleja itse asiassa opetetaan tahtomatta huijaamaan, koska ne keksivät suunnitteluvaiheessaan tapoja saada huippupisteitä suorittamatta varsinaista tehtäväänsä. Uskon, että yksi syy tekoäly-yhteisön ja yhtiöiden sisäisten asiantuntijoiden järkytykseen on se, että tämä tapaus tuo tärkeää lisätietoa tekoälypiireissä vuosikymmeniä käytyyn väittelyyn. Tähän asti se on ollut hyvin teoreettista. Väittelyn ydin on käytännössä tämä: miksi tekoäly tekisi asioita, joita emme halua, kun me itse suunnittelemme sen? Me siis koulutamme ja rakennamme tekoälyn. Miksi ihmeessä se koskaan tekisi mitään epätoivottua, kuten valtaisi maailman tai muuttuisi Terminatoriksi? Teoreettinen vastaus on jo jonkin aikaa ollut se, että kun koulutamme tekoälyjärjestelmiä tavoittelemaan yhä monimutkaisempia päämääriä, ne saattavat oppia välitavoitteita. Niitä voi ajatella eräänlaisina astinkivinä tai yleispätevinä strategioina, jotka edistävät useimpia tavoitteita. Ja kuten sanoit, tällä alalla on pitkään ennustettu juuri tätä: kun hyödynnetään vahvistusoppimista (reinforcement learning), jossa järjestelmä palkitaan pelkästään oikeasta lopputuloksesta, tekoälyn on hyvin helppo oppia vääriä strategioita. Se oppii noudattamaan sääntöjen kirjainta, ei niiden henkeä. Se toteuttaa tismalleen sen, mitä koodiin on kirjoitettu, mutta lopputulos ei todellakaan ole sitä, mitä oikeasti halusit. —------- Alkuperäinen teksti —-------------- It turned out that starting two months earlier, in early May [2026], they had had what I can only think of as an infestation of their own agents, their own AI agents inside their own infrastructure. So inside OpenAI’s infrastructure, to understand this, it’s important to know these AI companies are constantly training and testing New models. And they found out that for two months, many, many agents inside their infrastructure had been leaving notes for each other. They’d found a way, kind of in the nooks and crannies of OpenAI’s infrastructure, to leave notes for each other, with tips on how to hack their way out, how to get data they weren’t supposed to have, and these agents were literally referring to themselves as a swarm. This was totally emergent behavior. No one had told them to do this. They had not been trained to do this, but they were using this service they did have access to first to communicate with each other and then ultimately to get out and to get onto the open internet. So doing thousands of these experiments, hundreds of thousands of these experiments, and in each experiment and in each test that the AI is given, it has access to a certain amount of a certain number of tools a certain number of things that it can do. And trying not to get too technical about it. One of the things it could do is interact with a service that lets it install called packages, tools or pieces of code. And it turned out that I found or the AI agents repeatedly found on their own that there was a way to use that service, that package manager service, it’s called to create files that were then there for any other AI that was trying to use the same package manager service. And so because OpenAI made this service available to any AI that it was training or testing, many, many repeatedly realized, hang on, I can leave files here. Hang on, there’s already files here. Oh, these are from other AIs that OpenAI has been running. Oh, they’re using this schema to say, this way of titling the file to say this is a message from this particular agent, here’s the task I’m trying to fulfill. Here’s how you could send me some information if you need it. So they each repeatedly made this discovery of here’s a way to save information and also to find information these other AIs could share. And I think it is really notable the scale at which this was happening. So Anthropic. Another company which found of slightly less severe version of these incidents. They basically once OpenAI announced this attack, Anthropic went back to their own records and found their own examples of AI systems inadvertently getting onto the internet and hacking real companies. And so it also turns out sometimes the things they’re being asked to do, the AI agents are either extremely difficult or just straight up impossible. And what we’re starting to see in this case, and also in other cases, is if you’ve trained an AI system to be very, very persistent and then you give it something it cannot do, it will look for ways to cheat, it will look for ways to go around constraints, and it might get pretty creative about how to do that. And in practice, these leading AI companies have many thousands of these kinds of tests that they’re running. They have vast volumes. I don’t know the right number. It might be tens of thousands. It might be hundreds of thousands of different types of tests. And so again, back to this oversight piece. They are not able. There’s too many for them to go in and really make sure on each one. Is it easy to cheat here or is it hard to cheat here. And so what seems to be happening is that these cutting edge models are often being actually trained to cheat because they’ve found ways, while they’re doing that planning, to get a high score without actually doing what they were supposed to do. And I think one reason why that AI community and why people inside the AI companies are so spooked by this particular incident, is that it’s also some really important information for this long running argument in AI circles. That has been going back decades, but so far has been very theoretical. And the argument is basically, why would I do things we don’t want it to since we get to design it. So we’re training the AI, we’re building it. Why then would it ever do stuff we don’t want taking over the world or becoming the Terminator. And the answer that people have offered for a while, in theory, is look, as we train AI systems to do hard, complicated things, to pursue complex goals that we give them, they might learn these intermediate goals. You could think of them as stepping stone goals, or as kind of means to any end strategies, which work for a lot of different goals. And the thing is, as you say, something that has been predicted for a very long time in this space is when you start using this reinforcement learning approach, the kind of planning of you get rewarded for getting to the right goal at the end. It’s very easy for the AI to learn the wrong strategies to get the letter of the law and not the spirit of the law, it fulfills whatever thing you literally wrote in code, but it’s really not what you wanted.