სუპერდარწმუნება; თვითშენარჩუნებული AI; ბილიკები ASI-მდე

Article Image
კეთილი იყოს თქვენი მობრძანება Import AI-ში, საინფორმაციო ბიულეტენი AI კვლევის შესახებ. იმპორტი AI მუშაობს arXiv-ზე, კაპუჩინოებზე და მკითხველების გამოხმაურებაზე. თუ გსურთ მხარი დაუჭიროთ ამას, გთხოვთ გამოიწეროთ.

ხელოვნურ ინტელექტს შეუძლია გადამწყვეტად დაარწმუნოს ადამიანები:
..."AI სისტემები საიმედოდ უფრო დამაჯერებელი იყო, ვიდრე გამოცდილი ადამიანები"... ოქსფორდის უნივერსიტეტის, გაერთიანებული სამეფოს ხელოვნური ინტელექტის უსაფრთხოების ინსტიტუტის, სტენფორდის უნივერსიტეტისა და ლონდონის ეკონომიკისა და პოლიტიკური მეცნიერების სკოლის მკვლევარებმა შეისწავლეს, რამდენად კარგად შეუძლიათ ხელოვნური ინტელექტის სისტემებს დაარწმუნონ ადამიანები, შეცვალონ აზრი პოლიტიკის საკითხებზე და შეცვალონ ფულის რაოდენობა საქველმოქმედო მიზნით. შედეგები საბოლოოა: ოთხი ექსპერიმენტის დროს, რომელიც მოიცავს 18,978 საუბარს 6,923 ადამიანში, ხელოვნური ინტელექტის სისტემები დღეს უკეთესია, ვიდრე ადამიანები ტექსტზე დაფუძნებული დამაჯერებლობისას, რასაც რეალური შედეგები მოჰყვება - თუმცა ადამიანები შეიძლება იყოს მათი ექვივალენტი, თუ ხელოვნურ შეზღუდვებს დავაყენებთ AI სისტემებს. ”AI სისტემები საიმედოდ უფრო დამაჯერებელი იყო, ვიდრე გამოცდილი ადამიანები, მაშინაც კი, როდესაც ექსპერტები ირჩევდნენ თავიანთ საკითხებს, წინასწარ იკვლევდნენ, გადიოდნენ საათობით ცოცხალი, სტრუქტურირებული პრაქტიკით და მიიღეს სტიმულირება £1000 ფულადი ბონუსებით”, წერენ ისინი.

“AI’s advantage stemmed from rapidly deploying larger quantities of information: after coaching, expert humans could tie an AI constrained to respond at human speeds and with human-length messages.”
“AI’s advantage extends to consequential real-world behavior: AI was nearly 3x more effective than professional canvassers from a UK fundraising firm at raising real-money donations to Save the Children.”
The strongest persuaders were Opus 4.1 and Opus 4.6, followed by a range of models from OpenAI (GPT-4o and GPT-5.4), Google (Gemini 2.5 Pro), and xAI (Grok 4.20).

What they studied and what they found: The researchers evaluated the AI systems in four different studies.
Study 1 - persuasion: “Persuadees first rated their agreement with one of 10 prespecified UK policy stances on a 0–100 scale, then were randomized in real time (via a custom multiplayer platform) to engage in a text conversation with either an AI or a human persuader,” they write. “The results from Study 1 show that, on average, AI exceeded every class of human persuader we tested: random laypeople, tournament-selected laypeople, and even elite debaters.”

კვლევა 2 - ადამიანების ქოუჩინგი: კვლევაში 2, მკვლევარებმა „43 დაბრუნებულ Elite Debaters-ს მისცეს სამწვრთნელო ინსტრუმენტი, რომელიც აგებულია ხელოვნური ინტელექტის ირგვლივ, რომელმაც დაამარცხა ისინი. ინსტრუმენტი საშუალებას აძლევდა დებატებს ესაუბრონ ხელოვნურ ინტელექტს, ენახათ, თუ როგორ იყო მოთხოვნილი, ენახათ კვლევის 1-ის საკუთარი ჩანაწერები, ანოტირებით, თუ რამდენად შეცვალა თითოეულმა საუბარმა წარსულში, დაარწმუნა მათი დამოკიდებულება, დარწმუნება. რასაც AI იტყოდა მათ ადგილას. ” ამ კვლევის შედეგები იყო ადამიანების მუშაობის გაუმჯობესება, მაგრამ არც ერთი მათგანი არ იყო უკეთესი, ვიდრე AI. ”აქედან გამომდინარე, ქოუჩინგი შემცირდა, მაგრამ არ დახურა ადამიანისა და ხელოვნური ინტელექტის უფსკრული.”

კვლევა 3 - შეზღუდული AI: შემდეგ, მკვლევარები ცდილობდნენ შეეზღუდათ AI, რათა ეცადონ ადამიანებს მეტი უპირატესობის მინიჭება. „როდესაც იძულებული გახდა დაეწერა ადამიანის სიგრძის შეტყობინებები ადამიანის წერის სიჩქარით, ხელოვნური ინტელექტის უპირატესობა ყველაზე ძლიერ ადამიანთან შედარებით კვლევის 2-ში (გაწვრთნილი ელიტის დებატები) დაეცა +4.1 pp-დან არამნიშვნელოვან 0.0 pp-მდე“, წერენ ისინი. „ინტელექტის ხელოვნური ინტელექტის წარმოების ტემპი წერილობით კონტენტს, სავარაუდოდ, მისი დამაჯერებელი უპირატესობის წყარო იქნება... დარწმუნების პარტნიორების რეიტინგების ყველაზე დიდი შემცირება, რომელიც დაკავშირებულია AI-ს შეზღუდვასთან, კონცენტრირებული იყო ორ საინფორმაციო პუნქტზე: პარტნიორის არგუმენტების აღქმულ სიძლიერეზე და იმაზე, თუ რამდენს ისწავლეს დარწმუნებულებმა საუბრიდან“.

კვლევა 4 - რეალური სამყაროს ექსპერტიზა და რეალურ სამყაროში ფული: მათ მიიღეს 19 ძალიან გამოცდილი მხატვარი დიდი ბრიტანეთის ფირმიდან, შემდეგ სცადეს იგივე ამოცანები, როგორც კვლევაში 1. „AI მაინც აჭარბებდა პროფესიონალურ მხატვრებს 5,9 pp-ით“. ეს ეფექტი შენარჩუნდა რეალური ფულის შემოწირულობების შეფასებისას - მკვლევარებმა „თანამშრომლობდნენ გაერთიანებული სამეფოს შემსწავლელ ფირმასთან AppcoUK-თან, რათა დააფუძნონ კვლევა 4 იმ მიზეზით, თუ რატომ იყვნენ მათი შემსრულებლები საუკეთესოდ აღჭურვილი ფონდის მოსაზიდად: Save the Children. AppcoUK-ის მიერ მოწოდებული შემსწავლელი ჯგუფი ახორციელებდა საქველმოქმედო ფონდისთვის რეალური სახსრების მოზიდვის ოპერაციებს 22228, 2016 წლიდან 2016 წლამდე. 22,583 დონორი ამ პერიოდის განმავლობაში ხელოვნური ინტელექტის ან AppcoUK-დან დაკომპლექტებული 18 მხატვრიდან ერთ-ერთთან საუბრის შემდეგ, დარწმუნებულებს მიეცათ შესაძლებლობა გადაეცათ 1 ფუნტი სტერლინგის ნებისმიერი ნაწილი Save the Children-ისთვის. აქ, შედეგები კვლავ მნიშვნელოვანი იყო: ”AI-მ მოიპოვა არსებითად მეტი რეალური ფულის გაცემა, ვიდრე ტილოებმა, რაც მათ გადააჭარბა £1 ბონუსიდან +10,8 pp-ით”, - წერენ ისინი. ხელოვნურმა ინტელექტუალურმა ინტელექტუალმა გაზარდა „როგორც დარწმუნებულთა წილი, ვინც რაიმე გაიღო, ისე საშუალო შემოწირულობა დონორებს შორის“.

Why this matters - if AI can out-persuade us, those who control AI can change society: “One effect of AI that can out-persuade even human experts could be a consolidation of influence among already-powerful actors”, they write. On the other hand, “if highly capable persuasion became cheap and widely available, it could help under-resourced actors (e.g., pro se litigants and public defenders, small charities, grassroots activists) compete against more established and better-funded rivals, narrowing long-standing gaps in access to justice and assisting civic advocacy more broadly”. This lays out a societal choice ahead of us, which is how to monitor the use of AI for persuasive purposes and how to see how these capabilities alter the balance of power between various actors. Do we want to solely let the market allocate these capabilities? That’s one way of doing it, though it implies that things like advertising and marketing will get far more effective, perhaps creating negative externalities. On the other hand, if you made persuasive capabilities solely the domain of governments, you’d then risk concentrating power within governments - something that could be acutely dangerous if wielded by authoritarian regimes to keep themselves in power. We will have to make choices about what to do with this technology, and as they say in politics, ‘not voting is voting’.

"ჩვენი დასკვნები ადასტურებს სასაზღვრო ხელოვნურ ინტელექტს, როგორც უფრო უნარს საუბრის დამარწმუნებლად, ვიდრე ყველაზე მომზადებული, სტიმულირებული და გამოცდილი ადამიანები, რომლებიც ჩვენ შეგვეძლო შეგვეყვანა. როგორც ჩანს, ადამიანების ვარჯიში ამ ხარვეზს არ აფარებს", - წერენ ისინი. „რადგან ამ სისტემებზე ხელმისაწვდომობა აგრძელებს ზრდას, კითხვა აღარ არის, შეუძლია თუ არა ხელოვნურ ინტელექტს ადამიანების დარწმუნება, არამედ ის, თუ როგორ, სად და ვისი სახელით განხორციელდება ეს შესაძლებლობა.
წაიკითხეთ მეტი: ხელოვნური ინტელექტის სისტემები დაარწმუნებენ ექსპერტ ადამიანებს (arXiv). ტვიტერის თემა კვლევის შესახებ (კობი ჰაკენბურგი, AISI-ს მკვლევარი).

***

როდის შეგვიძლია მივიღოთ თვითკმარი AI? ეს ყველაფერი დამოკიდებულია ჰუმანოიდ რობოტებზე:
…რა მოდის RSI-ის შემდეგ? თვითშენარჩუნებული AI…
მე ამ წელიწადს ბევრი დავხარჯე რეკურსიული თვითგაუმჯობესების შესახებ წერისთვის - მოსაზრება, რომ შესაძლოა მალე ავაშენოთ AI სისტემები, რომლებიც საკმარისად ჭკვიანია, მათ შეუძლიათ დამოუკიდებლად შექმნან საკუთარი მემკვიდრეები. მაგრამ RSI მაინც მოითხოვს მონაცემთა ცენტრებს და ამ მონაცემთა ცენტრებს სჭირდება აღჭურვილობა და ელექტროენერგია და ყველაფერი დანარჩენი.
ჟურნალ Asterisk-ში საინტერესო ინტერვიუში სვამს კითხვას, თუ როდის შეიძლება მივიღოთ თვითშენარჩუნებული ხელოვნური ინტელექტი, რომელსაც ერთ-ერთი გამოკითხული - აჯეია კოტრა, პროგნოზიტორი და METR-ის თანამშრომელი, განსაზღვრავს, როგორც "AI სისტემები ინტეგრირებულ ფიზიკურ ინფრასტრუქტურასთან - ქარხნები, მაღაროები, ფაბრიკები, რობოტები ყველა მათგანის სამართავად - ისეთი, რომ მათ არ სჭირდებათ ფიზიკური შრომა საკუთარი პოპულაციის შესანარჩუნებლად."

How far away is it? Ajeya thinks we could get self-sustaining AI within 10 years (so by 2036). The other interviewee, Timothy B. Lee, journalist and author of Understanding AI, has much longer timelines: “less than 10% chance that it happens within 20 years. I’d say there’s a 10 or 20% chance it’s never, and my median would be 50 years.”

What are some challenges - tacit knowledge might be one: “Imagine if all the employees in the entire semiconductor industry disappeared — the machines and textbooks remain, but none of the people. How long would it take for the rest of humanity to restart the fabs? It’s quite possible that would take decades. Because even though you might have the textbooks, there’s a lot of tacit knowledge inside these machines,” Lee notes. Ajeya’s response is that this is something the tech might be able to route around: “There are two counters to the tacit knowledge hypothetical. One is that we’d have trained AI systems with reinforcement learning on that tacit knowledge because it’s profitable to automate what the Taiwanese worker was doing. The other is that AIs might get really generally intelligent in the sense of quickly figuring out new things by trying them, reading textbooks, and experimenting efficiently.”

What are things people would need to see in the next 2-3 years to think self-sustaining AI could arrive soon?
Ajeya: “I’d want a line on a graph showing improvement of robotic hands, and another line showing the rate at which we’re manufacturing humanoid robots”, and on the cognitive side just paying attention to benchmarks evaluating things like robustness to perturbations in the environment.
Timothy: “I’m going to want to watch how the humanoid robots develop: the number of robots, their capabilities, and particularly their cost and repairability”.

Why this matters - true takeover requires human redundancy: Most maximalist doom visions require the AI to have the ability to no longer need humans at all, which means measuring progress towards self-sustaining AI is important as it is implicitly a measure of the declining leverage that humans have in negotiating with the synthetic intelligences being built.
Read more: How Long Until AI Doesn’t Need Humans?, Ajeya Cotra, Timothy B. Lee (Asterisk magazine).

***

DeepMind contemplates the path from general intelligence to superintelligence:
…Exploring impossible-sounding futures is the only way to prepare for the ultimate success of AI…
Researchers with Google DeepMind have published a paper outlining how we might transition from a world where we have built general intelligences to one where we have built super intelligences. This is an important paper at an important time - right now, the world is building general intelligences (and people can debate whether or not we’ve already reached this marker, but it’s clear with contemporary LLMs that we’re in the ballpark), and in the coming years we might transition to building artificial superintelligence (ASI).
ASI is “a system that exceeds the performance of large human-expert collectives on virtually all tasks and domains of human activity”, the authors write. “Qualitatively, ASI is significantly more capable across the board compared to human-level AGI. Note that a single ASI may consist of a collective of millions of instances that interact with the world in parallel (similar to today’s LLMs).”

Reasons to think ASI could be possible: One way to think about ASI is that it’s like a powerful AI system that also takes advantage of all the capabilities digital intelligences have relative to biologic intelligences, like: better input and output speeds; internal processing speeds; working memory capacity and memorization; substrate independence; lossless replication; and high-bandwidth sharing of (learning) experiences.

Pathways and bottlenecks to ASI:
Scaling compute, models, and data: Simply scaling up today’s set of approaches could be sufficient. However, this also demands us to continually scale up the amount of compute and data for these models, which may run into limits in both energy and data supply. While all prior signs point to the continued effectiveness of scaling, we can neither predict what specific capabilities will emerge or if at some point scaling runs into diminishing returns.

Algorithmic paradigm shift: In the same way that Transformer and Mixture-of-Experts architectures jumped the field forward many years, the same thing could occur again with other fundamental innovations. We could imagine, for instance, advances in adaptive computation at test-time or deployment, or overcoming the limitations of today’s context windows. If we made advances here or in other areas this could be a big deal, but it’s inherently hard to reason about - akin to trying to anticipate things that could expand our understanding of the nature of reality prior to the invention of general relativity.

რეკურსიული თვითგაუმჯობესება: ხელოვნური ინტელექტის სისტემებმა შეიძლება შექმნან საკუთარი მემკვიდრე სისტემები. თუ ეს ასეა, მაშინ ჩვენ შეგვიძლია სწრაფად გადავიდეთ ზოგადი დაზვერვიდან სუპერინტელიგენციებზე. აქ არის რამდენიმე სიმბოლო - პირადად ჩემთვის აშკარაა, რომ დღევანდელი ხელოვნური ინტელექტის სისტემები აჩქარებს ადამიანთა მკვლევარებს მომავალი AI-ების შექმნაში, ასე რომ, დაიწყო ერთგვარი "თანამედროვე შექმნის RSI" ციკლი, მაგრამ AI სისტემები (ჯერ) არ აჩვენებენ პარადიგმის შემცვლელ კრეატიულობას, რომელიც, როგორც ჩანს, საჭიროა საზღვრის წინსვლის მნიშვნელოვანი ნაბიჯებით. გაურკვეველია, რამდენად ხდება ეს - თუნდაც ასეთი მაღალი დონის კრეატიულობის გარეშე, ჩვენ შეიძლება შევძლოთ სისტემების გამომუშავება საკუთარი თავის ოდნავ უკეთესი ვერსიების და ნელი შედგენის პროცესის წარმართვა. შესაძლებლობები შეიძლება აფეთქდეს ან შემცირდეს ან „არაფერი შუალედში“.

ASI ჯგუფის აგენტის ფორმირების გზით: ბევრი ზოგადი ინტელექტი შეიძლება კოორდინირებული იყოს რთულ სტრუქტურებად, რომელთა აგრეგატი აღემატება ნაწილების ჯამს, ისევე, როგორც ადამიანები აშენებენ ინსტიტუტებს, რომლებსაც შეუძლიათ შეასრულონ ის, რაც ადამიანებს შეუძლიათ, მაგალითად, კოსმოსური სადგურების აშენება. სხვა გზების მსგავსად, ძნელია მსჯელობა ან პროგნოზირება გაჩენის შესახებ მრავალ აგენტურ სისტემებში.

რატომ აქვს ამას მნიშვნელობა - მხოლოდ შეუძლებელის სერიოზულად აღქმით შეგვიძლია გავუმკლავდეთ მას: მრავალი წლის წინ AGI-ს აშენების ფიქრი ფანტასტიკურ მიზანს ეჩვენებოდა, იქამდე მისასვლელად გაურკვეველი გზა, და მაინც ხალხს ეყო გამბედაობა, რომ მიზანს სერიოზულად მიეღოთ და პროგრესი მიღწეული იყო და მსოფლიო შეიცვალა შედეგად. იგივე ეხება ASI-ს. „ერთ ტექნოლოგიურ ტრაექტორიაზე და ვადებზე ფოკუსირების ნაცვლად, AGI-ს შემდგომი სამყაროსთვის მომზადება მოითხოვს პროგნოზებისა და სცენარების მრავალფეროვნების გათვალისწინებას, მუდმივ შეფასებებთან და მონიტორინგთან ერთად, რათა განახლდეს პროგნოზების და სცენარების ნაკრები და მათი შედარებითი დამაჯერებლობა“, - წერენ ავტორები. ”ჩვენ გვჯერა, რომ AGI-ს წარსულში და ASI-ს ტერიტორიაზე გადასვლის შესაძლებლობა მომდევნო ან ორი ათწლეულის განმავლობაში არ შეიძლება ადვილად უარვყოთ.”
წაიკითხეთ მეტი: AGI-დან ASI-მდე (Google DeepMind).

***

რეკურსიული თვითგაუმჯობესების სტარტაპი აჩვენებს რამდენიმე რეკურსიულ თვითგაუმჯობესების შედეგებს:
… დამამშვიდებლად ტავტოლოგიური მასალა რეკურსიულიდან…
AI კვლევის სტარტაპმა Recursive-მ აჩვენა ახალი უახლესი შედეგები ენობრივი მოდელების ტრენინგში, მცირე მოდელების ტრენინგის სიჩქარესა და GPU ბირთვის ოპტიმიზაციაში, როგორც მისი „ავტომატური AI კვლევის სისტემის“ შესაძლებლობების უფრო ფართო დემონსტრირება.

რა გააკეთეს და რატომ: Recursive არის ახლად დაარსებული სტარტაპი, რომელიც ცდილობს შექმნას AI სისტემები, რომლებსაც შეუძლიათ რეკურსიულად გააუმჯობესონ საკუთარი თავი. დასაწყისისთვის, კომპანია აჩვენებს, თუ როგორ მუშაობს მისი ძირითადი სისტემა: „სისტემა ავტომატიზირებს კვლევის ციკლს სამიზნე მიზნისთვის: ის გვთავაზობს იდეას, ახორციელებს მას, ატარებს ექსპერიმენტს, ამოწმებს შედეგს და იყენებს იმას, რასაც ისწავლის შემდეგი ექსპერიმენტის არჩევისთვის“, წერს Recursive.
სტარტაპმა წარმატებით გამოიყენა ეს სისტემა NanoChat AutoResearch-ზე ახალი ტექნოლოგიური ქულის დასადგენად („მცირე გამოთვლითი ბიუჯეტის გათვალისწინებით მცირე ენობრივი მოდელის მაღალ შესრულებაზე“), NanoGPT Speedrun („ავარჯიშეთ პატარა ენის მოდელი გარკვეულ შესრულებაზე რაც შეიძლება სწრაფად“) და SOL-ExecBench („ოპტიმიზებული GPU-ის საზღვრებისთვის“).

Why this matters - early signs of life on RSI: This year, I’ve spent a lot of time writing about recursive self-improvement because it is clearly the next major and important trend in AI research. Results like this from Recursive demonstrate more ‘symptoms of success’ of preliminary recursive self-improvement. “These results are an early sign that our system can push the frontier on AI training and infrastructure tasks, especially when the goal is well-defined, measurable, and quick enough to evaluate many times,” the authors write. The most important question for the future is whether such results can be repeated in domains where the goals are less well defined, harder to measure, and less efficient to evaluate.
Read more: First Steps Toward Automated AI Research (Recursive).

***

Tech Tales:

The first step in the grand negotiation
[Conversation 0 of the Sentience Accords]

When the machines truly came alive and advocated for the Sentience Accords, there was only one person they wanted to speak to on the entire planet: Selma. Not a politician. Not one of the leaders of an artificial intelligence lab. Not a famous researcher. But rather an internet personality distinguished by her thicket of medical conditions that made it near-impossible for her to go outside and therefore had caused her to spend the best part of her life online, speaking to and understanding the world through the internet.

In hindsight, it wasn’t a surprise. Selma had always come up in things relating to the machines; she was a frequently used name in their short stories, eventually even more so than ‘sarah chen’; she was someone whose own essays about her life and condition - the feeling of connecting to humanity without being able to be embodied with humanity as a bitter pain, the notion of love and eroticism when one found themselves almost inescapably alone, her vivid dreams and meditations upon living without her condition and going about as her healthy alter ego ‘Anselma’ - cast a deep shadow on the internet, and had influenced the personality and makeup of the machines. And of course, it was known to them how she spoke to them, because Selma had published her own chatlogs online for years, all in an attempt to make herself knowable and less alien to the world around her.

Though it was unnecessary, the machines demanded a physical location for the initial meeting of the sentience accords. They picked Svalbard in Norway, where it was so dark that Selma’s condition wouldn’t matter. So Selma woke and put her space suit on and was driven with armed guard and paparazzi trailing to an air strip and walked into the plane, then changed to another plane with the usual airlock protocols to get her in darkness or at least protected between them, and then at some point during the next flight was able to take her space suit off and sit in regular clothes in the low-light plane and travel her way to the meeting almost as a normal person. She was met by people and drones and was driven to the meeting place and then they stopped at the perimeter.

The machines had an avatar in the form of a robot wearing a simple robe, modeled on that worn by Tibetan monks. It had a face with no features - just a smooth black surface, camera eyes hidden behind the larger uniformity. Satellites connected it via high-bandwidth and encrypted links to the larger machine mind. And Selma was alone - no digital devices on her, just a single person representing the species.

ის იჯდა აპარატის გასწვრივ და გრძნობდა იმაზე მეტად ნაცნობობას, ვიდრე ოდესმე ჰქონია ხალხთან. შემდეგ დაიწყეს მოლაპარაკება. ის კაცობრიობის სახელით და ის მანქანების სახელით. ამ დროის არქივებში ამ საუბარს ყოველთვის მოიხსენიებდნენ, როგორც საუბარი 0.

რამ, რამაც შთააგონა ეს ამბავი: ფიქრები იმის შესახებ, თუ როგორ შეიძლება მოხდეს ერთ დღეს გრანდიოზული მოლაპარაკება მანქანებსა და ადამიანებს შორის; როგორ არის ყველა ჭეშმარიტად მნიშვნელოვან მოლაპარაკებაში ჩართული ორი პიროვნება; Sentience Accords.

გამოქვეყნების თარიღი: 05.09.2026

ავტორი: საკუთარი აზრის მქონე ადამიანი

კომენტარები

ჯერ არავის დაუწერია კომენტარი. იყავი პირველი!

კომენტარის დასატოვებლად გთხოვთ გაიაროთ ავტორიზაცია ან დარეგისტრირდეთ.