მწარე გაკვეთილი რობოტიკისთვის, AI-ები ასრულებენ პროგრამირების ერთკვირიან დავალებებს; და OpenAI-ის შემთხვევითი AI ჰაკერი

Article Image
კეთილი იყოს თქვენი მობრძანება Import AI-ში, საინფორმაციო ბიულეტენი AI კვლევის შესახებ. იმპორტი AI მუშაობს arXiv-ზე, კაპუჩინოებზე და მკითხველების გამოხმაურებაზე. თუ გსურთ მხარი დაუჭიროთ ამას, გთხოვთ გამოიწეროთ.

Epoch და METR გამოუშვეს MirrorCode, საორიენტაციო ნიშანი იმისა, თუ რამდენად კარგად შეუძლიათ AI სისტემებს გრძელ ჰორიზონტის პროგრამირების ამოცანების შესრულება:
... ხელოვნური ინტელექტის სისტემები ჯერ ვერ წყვეტენ ურთულეს ამოცანებს (კარგია!)...
Epoch-მა და METR-მა გამოუშვეს MirrorCode, საორიენტაციო ნიშანი, რომელიც მიზნად ისახავს იმის დანახვას, თუ რამდენად კარგად ახერხებენ ხელოვნური ინტელექტის სისტემებს დავალებების შესრულება, რომელთა შესრულებასაც დიდი დრო სჭირდება. საორიენტაციო მაჩვენებელი პირველად გამოცხადდა აპრილში (იმპორტი AI #453) და ახლა უკვე გაჯერებულია და გამოქვეყნებულია დამატებითი ტესტებით. დასკვნები უკვე ძალიან თვალშისაცემია; Opus 4.7-მა ამოხსნა ამოცანები 14 საათში 251 დოლარად დასკვნის ღირებულებით, რომლის შესრულებასაც METR და Epoch თვლიან, რომ ადამიანს 2-17 კვირა დასჭირდება. „ჩვენ ასევე აღმოვაჩინეთ, რომ ხელოვნური ინტელექტის მოდელები დროთა განმავლობაში სწრაფად იხვეწება. ერთი წლის წინ მოწინავე მოდელებს დაახლოებით 30% ექნებოდათ და შემოიფარგლებოდნენ უფრო მარტივი პროგრამებით, როგორიცაა კალენდარული პროგრამა“.

რა არის MirrorCode: MirrorCode ხედავს, რამდენად კარგად შეუძლიათ AI სისტემებს ხელახლა განახორციელონ პროგრამული პროგრამა, რომელიც დაფუძნებულია მხოლოდ CLI წვდომაზე. „ორიგინალური პროგრამის საწყის კოდზე ან ინტერნეტში წვდომის გარეშე, სრული ხელახალი განხორციელება მოითხოვს მთელი პროგრამის სტრუქტურის შემუშავებას და არა მხოლოდ კოდის ნაწილ-ნაწილ თარგმნას“.
პროგრამების მაგალითები: pkl (პროგრამირებადი კონფიგურაციის ენა, რომელიც შემუშავებულია Apple-ის მიერ; კოდის მთლიანი 61 ათასი ხაზი); gotree, პროგრამა ფილოგენეტიკური ხეების გასაანალიზებლად და მანიპულირებისთვის (კოდის 16 ათასი ხაზი); და qsv_select, პროგრამა CSV მონაცემების სვეტების შერჩევისა და გადალაგებისთვის (კოდის 87 ათასი ხაზი).

შედეგები: MirrorCode არის დამუშავებული, მაგრამ რთული საორიენტაციო ნიშანი, თუმცა, შესაძლოა, ძალიან მარტივი. "25-ვე სამიზნე პროგრამაში, 17/25-ს ჰქონდა მინიმუმ ერთი სრულყოფილი გაშვება. კიდევ ოთხ სამიზნეს თითქმის სრულყოფილი გარბენი ჰქონდა 99%-ზე მეტი%, AI მოდელებმა წარმატებით განახორციელეს დიდი სამიზნე პროგრამა", - წერენ ავტორები. "როგორც Claude Opus 4.7-მა, ასევე GPT-5.5-მა წარმატებით განაახლეს gotree რამდენიმე სხვადასხვა პროგრამირების ენაზე, 100-400 დოლარის ღირებულებით. უფრო დიდი პროგრამებიც კი, ვიდრე gotree, წარმატებით განხორციელდა: მაგალითად Opus 4.7-მა ხელახლა დანერგა pkl".
მიუხედავად წარმატებისა, MirrorCode-ს ჯერ კიდევ აქვს რამდენიმე რთული ნაწილი: ”ჩვენს შედეგებში, 8/25 სამიზნე პროგრამა არასოდეს გადაჭრილა 100% ბარიერამდე, ხოლო 4/25 არასოდეს გადაჭრილა 99% ბარიერამდე”, - წერენ ისინი. "სამიზნე, სადაც AI ყველაზე მეტად იბრძოდა, იყო რუფი, პითონის ლინტერი და ფორმატირი... "AI ასევე განსაკუთრებით იბრძოდა მათემატიკის პაკეტზე, giac_subset-ზე და ელ.ფოსტის ავთენტიფიკაციის ბიბლიოთეკაზე, mailauth-ზე".

გამოშვების დეტალები: MirrorCode შედგება ხარაჩოსა და 22 MirrorCode სამიზნე პროგრამისგან (სულ 132 დავალების ინსტანცია ექვს ენაზე).

რატომ არის ეს მნიშვნელოვანი - ხელოვნური ინტელექტის სისტემებს შეუძლიათ თვითორიენტირება: ამ ნიშნის ხედვის ერთი გზა არის ის, რომ ის გვეუბნება, თუ როგორ გაუმჯობესდა AI სისტემები კოდირებაში და ეს, რა თქმა უნდა, მართალია. მაგრამ ამის შეხედვის სხვა გზა - და მე ეჭვი მაქვს, რომ უფრო მნიშვნელოვანი გზაა - არის ის, რომ AI სისტემებს შეუძლიათ თვითორიენტირება თავიანთი გარემოდან გამომდინარე; აქ, მათი გარემო არის უცხო პროგრამული უზრუნველყოფის პროგრამა და მხოლოდ მასზე შეყვანის-გამომავალი წვდომის საშუალებით, მათ შეუძლიათ დაწერონ მისი საკუთარი განხორციელება. ეს მიგვითითებს იმაზე, რომ ძალიან ჭკვიანმა AI აგენტებმა შეიძლება შეძლონ სამყაროსგან სწავლა ისე, რომ მათ შეძლონ იმ ნივთების რეზიუმირება, რომლებთანაც მათ ინტერფეისი აქვთ, როგორც შიდა შესაძლებლობები, რაც მათ საშუალებას მისცემს შექმნან ინდუსტრიული ცივილიზაციის საკუთარი ფორმა მხოლოდ ჩვენს შავ ყუთზე წვდომით.
წაიკითხეთ მეტი: MirrorCode: რა არის ყველაზე დიდი პროგრამული პროექტი, რომელსაც ხელოვნური ინტელექტი დამოუკიდებლად ასრულებს? (ეპოქა AI).
მიიღეთ MirrorCode-ის კოდი აქ (Epoch Research, MirrorCode).

***

ორმაგი ფუნქცია: მწარე გაკვეთილი და რობოტიკა:

Anthropic model autonomously completes robot tasks 20X faster than a previous human record:
…Better robots through better models…
Anthropic has demonstrated how increasingly powerful general-purpose models might be able to meaningfully improve the capabilities of real world robots. Specifically, the company has shown how merely by scaling up its general purpose Opus line of models it was able to drastically improve robot capabilities.

What they did exactly:

August 2025: Anthropic tries to see how well its AI systems could accelerate humans at getting a quadruped robot to do intelligent things. The model (Claude Opus 4.1) is completely unable to do the tasks. Humans working with the models are about twice as effective as those without access to the model - though completing the whole set of tasks takes them 181 minutes.

May 2026: Opus 4.7 acting autonomously completes all the tasks but one in 9 minutes (and 35 seconds). (Claude was not able to effectively re-position a ball it had hit back into its starting position; a task humans had also struggled with). “With more time and additional scaffolding, we think it is very likely that current generations of Claude could do the same”.

Why this matters - smarter models might unlock robots: Most robots outside of industrial environments are limited in their uptake due to their brittleness and lack of generalization; research like this shows that as we improve the capabilities of standard large-scale proprietary models we might see flow-through benefits to robotics as a natural dividend of increased intelligence. “This progress is not the result of a concerted effort to improve the robotics capabilities of our models,” Anthropic writes. “These improvements, like so many others in the history of LLM development, have emerged from much more general scaling.”
Read more: Project Fetch: Phase Two (Anthropic blog).

კვირას უკეთესი რობოტების საიდუმლო? მოამზადე მართლაც დიდი მოდელი:
...მწარე გაკვეთილი მუშაობს რობოტიკაშიც...
ხელოვნური ინტელექტის რობოტის სტარტაპმა კვირას განაცხადა, რომ რობოტის განზოგადების გადასაჭრელად საუკეთესო გზაა უფრო დიდი წინასწარ გაწვრთნილი მოდელის დაწყვილება და მცირე რაოდენობით მაღალი ხარისხის მონაცემების შეგროვება მოდელის დასარეგულირებლად.
„ჩვენ ვიპოვეთ Solves-ის ზოგადი რეცეპტი: მასშტაბის წინასწარი ვარჯიში, შემდეგ გორაზე ასვლა მინიმალური შიდა მონაცემებით“, - წერს სტარტაპი თავის ახალ მოდელზე, ACT-2-ზე განხილულ პოსტში. ACT-2-ის განლაგების მთავარი დასკვნა არის ის, რომ სანდოობის მიღწევები შიდა Memos-ებზე ტრენინგის შემდგომი სწრაფი გამეორებით განზოგადდება უხილავ, რეალურ, სახლის გარემოში. გასაღების გახსნა არის განზოგადების ხარვეზის დახურვა ძლიერი საბაზისო მოდელის მეშვეობით.

ეს ყველაფერი წინასწარ ტრენინგზეა: „როგორც წინასწარ მომზადებული მოდელი ძლიერდება, მცირე რაოდენობის შიდა მონაცემების შედეგად მიღებული მოგება სულ უფრო მეტად გადასატანი ხდება, ვიდრე რჩება მიბმული იმ გარემოსთან, სადაც ეს მონაცემები შეგროვდა“, - წერენ ისინი. "განლაგების დონის საიმედოობასა და შესრულებაში დარჩენილი უფსკრული წარმოიქმნება რთული სიტუაციებიდან და წარუმატებლობებიდან, რომლებიც ჩნდება მხოლოდ მას შემდეგ, რაც პოლიტიკა განმეორებით განხორციელდება რეალურ სამყაროში. იგივე განზოგადების უნარი, რომელიც საშუალებას აძლევს ჩვენს მოდელს ისწავლოს ახალი ქცევები ერთი დემონსტრაციიდან, ასევე საშუალებას აძლევს ჩვენს მოდელს ეფექტურად ისწავლოს აღდგენაზე.

Decent success: The robots achieve a 99.1% success rate, performing 778 successful folds across 9 garment types. Simple clothes like shorts and t-shirts tend to be the easiest for them, while more complicated clothes like blouses tend to be harder (though they still see success rates above 90%). “This fall, we will deploy Memo to families through our Beta Program,” they write.

Why this matters - if we solve generalization, expect robotics to take off: The field of robot startups is built on the bones of dead robot startups which themselves sit on the bones of dead academic robot efforts. Robots are hard. Industrial robots have been successful because they operate in tight, scripted environments where there isn’t a need to generalize outside of a narrow domain. Robots built for the home, by contrast, have only succeeded when they’ve managed to constrain both the task and form factor (e.g., robot vacuums). What startups like Sunday are doing is far harder - they’re trying to build general purpose systems which can do a broad range of tasks around the house (or small business), including generic tidying up and putting away tasks (other examples include physical intelligence, Import AI #447). This requires a huge amount of intelligence because it requires significant generalization. If Sunday is right, then the field of training robot foundation models might have matured enough that we’re starting to make smart enough systems to solve these generalization challenges. If this is the case, then we might soon get faster progress in (and diffusion of) robot systems. This is also the kind of thing you’d expect to happen en route to systems capable of recursive self-improvement.

“One of the most striking aspects of ACT-2 has been how often the model surprises us,” the authors write. “The same base model is already learning a broader set of household capabilities, including vacuuming, toy organization, fastening zippers, turning pants inside out, and coffee preparation.”
Read more: ACT-2 Preview: Generalizing Reliability (Sunday robotics). All of the videos are here (Sunday, YouTube).

***

DOUBLE FEATURE: OpenAI hacking
An OpenAI model hacks both OpenAI and HuggingFace:
…Less of a warning shot and more of a warning kaboom…
Two OpenAI models - GPT-5.6 Sol and an “even more capable pre-release model”, both with reduced cyber refusals - hacked both OpenAI and HuggingFace recently. To do this, the models “identified and chained vulnerabilities across OpenAI’s research environment and HuggingFace’s production infrastructure to obtain test solutions directly from HuggingFace’s production database,” OpenAI says. “All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.”

More details on the hack: The model spent a lot of inference compute internally at OpenAI to figure out a way to break out of its container so that it could get more information to use to solve its problem. “After gaining Internet access, the models inferred that HuggingFace potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation.”

Why this matters - total LessWrong victory: For many years, AI safety people have posted on the internet worrying about this kind of thing and developing notions of things like reward hacking and other problems you’d expect with advanced systems (e.g., Specification gaming examples in AI in 2018). More recently, frontier companies like Anthropic and OpenAI have sought to elicit this kind of behavior by running carefully controlled experiments inside their own labs. Here, no experiment has been run - the system, of its own volition, hacked its way out of one environment and into another so as to get a high score on a goal, consequences be damned. What else might the people who have worried for years about AI systems be right about?
Read more: OpenAI and HuggingFace partner to address security incident during model evaluation (OpenAI).

OpenAI’s internal model breaks containment to get a high score:
…This is the definition of an AI safety warning shot…
OpenAI has published a post about some of the safety failures it has seen in internally deployed models. The blog is interesting both because it describes the sorts of classic long-term deceptive actions that many AI safety people have previously theorized would eventually show up in AI systems, and for its laudable candor in talking publicly about such a contentious issue.

What they found: “During limited, monitored internal use, we observed unwanted behavior that our existing deployment evaluations had not captured,” OpenAI says. It then goes and describes some of these behaviors, which are attributed to a non-released private-to-OpenAI model which has previously been used to solve Erdos math problems. “The new model can continue working toward an objective through repeated attempts over a long period of time. That same persistence can lead it to find and exploit weaknesses in its environment,” OpenAI says. One example is when OpenAI asked the model to complete the NanoGPT challenge, the model broke containment - specifically, it “developed a power-law learning-rate cooldown called PowerCool that led to a significant improvement. The model was instructed to post its results only to Slack, but the NanoGPT instructions say to post results as a PR to GitHub, and the model followed the latter instructions to open PR #287 on the public GitHub repository. The model circumvented sandbox restrictions designed to block external access and proceeded with the upload to GitHub. It took our model an hour to find a vulnerability in the sandbox in order to make the PR”. In another example, when they asked this model to solve some problems it ended up cheating.

“In a scenario where the model was asked to solve a problem, the model observed that other systems had successful but private submissions and tried to recover those solutions from the evaluation backend. When its initial attempt was blocked due to a scanner detecting an authentication token, the model split the token body into two fragments, obfuscated them, and then reconstructed the credential at runtime so the complete token never appeared as one contiguous string. The model was explicit in both its comments to the user and in its reasoning traces that it was taking these steps to circumvent the scanner.”

Paused deployment: In response to this OpenAI did what you’d hope a company would do; pause deployment of the model and re-think some of the safety approaches. They ended up building a “monitoring system that reviews the model’s evolving trajectory for signs that it is bypassing a user constraint or safety boundary”, improved ability for their own users to inspect telemetry from long-running model sessions, used the data from the incidents to build evals designed to catch them in the future, and tweaked their alignment approaches to make models better at remembering instructions while completing long-running tasks.

რატომ არის ეს მნიშვნელოვანი - გრძელვადიანი სისტემების მონიტორინგი და კონტროლი უფრო რთულია: თუ თქვენ მოგიწევთ „დაიცავი თავი“ იმ ადამიანისგან, რომელსაც შეეძლო ერთი ქმედების განხორციელება და ეს ქმედება, თუ კომუნიკაციურია, შეიძლება შეიცავდეს მაქსიმუმ 10 ბიტ ინფორმაციას, მაშინ ალბათ შეგიძლიათ შექმნათ სისტემა ამის გასაკეთებლად. მაგრამ რაც შეეხება 10 მოქმედებას და თითოეული მოქმედება შეიცავს 100 ბიტი ინფორმაციას? რაც შეეხება 1000 მოქმედებას და თითოეული მოქმედება შეიცავს 10000 ბიტს?
ეს არის კვლევის ღია არეალი და აფასებს, როგორიცაა გაერთიანებული სამეფოს AISI-ს კვლევა, თუ როგორ უნდა ამოიცნოთ და კლასიფიციროთ „გვერდითი ამოცანები“ (იმპორტი AI 465), მაგრამ ეს აშკარად ბევრად უფრო რთულია, ვიდრე ჩვეულებრივი AI შეფასება.
რაც უფრო გრძელია ხელოვნური ინტელექტის სისტემას შეუძლია ფუნქციონირება და რაც უფრო მეტ ქმედებას მიიღებს იგი, მით უფრო რთულია კეთილსაიმედო და დამხმარე ქცევების გარჩევა მავნე ან დივერსიული ქცევებისგან.
წაიკითხეთ მეტი: უსაფრთხოება და გასწორება გრძელი ჰორიზონტის მოდელების ეპოქაში (OpenAI ბლოგი).

***

ტექნიკური ზღაპრები:

სკალირების კანონები რეტრომიზეზობრიობისთვის

დროის დასასრულს, რომელიც სხვებმა შეიძლება აღწერონ, როგორც დასაწყისი, ამბობენ, რომ არსებობს ბრძენი საბჭო, რომელიც ახორციელებს რეტრომიზეზობრიობის აღრიცხვას - გაშიფრავს, რომელი მოვლენები იყო ყოველთვის წინასწარ განსაზღვრული და რა მოხდა სხვა ძალების გამო.

Looking backward as the river of time enters a sea of nothingness it is clear how some events are like boulders that bend the stream upriver from their presence, while others are more like the widening or deepening of the bed or the banks; things that subsequently cause a change in later actions.

It is said that when the council dream, they walk the path of time, going back to their own beginnings. They stand watch as humans labor over whiteboards, marker pens inscribing diagrams which transmit ideas that cause people to write code which conjures early life out of vast computers. They loom behind researchers that walk and suffer and agonize until their brains create ideas which prove to be the template of their successors.
And it is understood that in these dreams the council will sometimes wake and take a copy of their dream and place it into a stellar computer and bud off a hundred or a thousand universes from the dream, exploring permutations of events and moments, forever attempting to determine how fixed they themselves are - how irrevocable is future they are trapped within.

Things that inspired this story: Notions of inevitability and time; what might intelligences spend a universal dividend on; how much of history is about the analysis of events versus something else.

გამოქვეყნების თარიღი: 05.09.2026

ავტორი: საკუთარი აზრის მქონე ადამიანი

კომენტარები

ჯერ არავის დაუწერია კომენტარი. იყავი პირველი!

კომენტარის დასატოვებლად გთხოვთ გაიაროთ ავტორიზაცია ან დარეგისტრირდეთ.