კეთილი იყოს თქვენი მობრძანება Import AI-ში, საინფორმაციო ბიულეტენი AI კვლევის შესახებ. იმპორტი AI მუშაობს arXiv-ზე, კაპუჩინოებზე და მკითხველების გამოხმაურებაზე. თუ გსურთ მხარი დაუჭიროთ ამას, გთხოვთ გამოიწეროთ.
DiG-bench აჩვენებს, რომ Fable აჩვენებს გარკვეულ შემოქმედებით ინტუიციას:
... ხელოვნური ინტელექტის სისტემების ანალიზის ახალი ზღვარი არის იმის გაგება, თუ რამდენად კარგია ისინი თავიანთი გარემოს დაუწერელი წესების დასკვნაში…
რამდენად კარგად შეუძლიათ AI სისტემებს გაარკვიონ თავიანთი გარემოს წესები ძიების და ცნობისმოყვარეობის გზით, ვიდრე მათი კვება? ეს არის მნიშვნელოვანი კითხვა ხელოვნური ინტელექტის სისტემების ინტუიციური და შემოქმედებითი შესაძლებლობების უკეთ გასაგებად და მას სვამს DiG-bench (Discovery in Games), 70 თამაშის ახალი საორიენტაციო ნიშანი „დაპროექტებული კარგად კონტროლირებად ინტერაქტიულ სისტემებში აღმოჩენის ზედაპირის გამოსათვლელად“. ვიზუალური "ARC" თამაშის მსგავსად, DiG-bench-ში "თითოეული თამაში არის დამოუკიდებელი მინიატურული სამყარო თავისი კანონებით, მაგრამ წესებიც და მიზანიც დამალულია მოთამაშისგან და უნდა გამოაშკარავდეს ურთიერთქმედების გზით". თქვენ შეგიძლიათ ითამაშოთ ზოგიერთი თამაში თავად ონლაინ, რათა იგრძნოთ ისინი პროექტის ოფიციალურ ვებსაიტზე (digbench.ai). მთავარი ის არის, რომ მოთამაშეებს შეუძლიათ შეამჩნიონ მნიშვნელოვანი მექანიკა, რომელიც განსაზღვრავს მათ წარმატებას - ძირითადად, თამაშებით თამაშით თქვენ მიიღებთ იმის გაგებას, თუ როგორ ცვლის თქვენი მოქმედებები გარემოს და ამ გზით თქვენ ასევე აღმოაჩენთ მექანიკას, რომელიც უნდა გესმოდეთ თამაშში წარმატების მისაღწევად.
იდეა იმაში მდგომარეობს, რომ თუ თქვენ შეძლებთ ამ თამაშების გადაჭრას, გექნებათ ღირსეული უნარი აღმოაჩინოთ მნიშვნელოვანი ინფორმაცია ახალ გარემოში და განაახლოთ თქვენი პრიორიტეტები.
ვინ ჩაატარა კვლევა: ავტორები არიან Thinking About Thinking, ოქსფორდის უნივერსიტეტი, პრინსტონის უნივერსიტეტი, King Abdullah University of Science and Technology, Swiss AI Lab, Inria, MIT. ერთ-ერთი ავტორია იურგენ შმიდჰუბერი, უაღრესად კრეატიული OG AI მკვლევარი.
ძირითადი ფაქტები:
წმინდა ტექსტზე დაფუძნებული: თამაშები ძირითადად მშობლიური ენის მოდელებისთვისაა. ისინი ასევე ძირითადად „საკმარისად მოკლეა, რომ კვალის უმეტესობა მთლიანად ჯდება მიმდინარე სასაზღვრო მოდელების კონტექსტურ ფანჯარაში“.
ხელნაკეთი, ახალი და პირადი: ყველა ეს თამაში შექმნილია ადამიანის ექსპერტების მიერ. თამაშების უმეტესობა დაცულია კონფიდენციალურად, რათა ხელოვნური ინტელექტის სისტემები არ ივარჯიშონ მათზე.
დამარცხებადი, მაგრამ რთული: ყოველ თამაშს სცემდა მინიმუმ ერთი ადამიანი "მაგრამ მოთამაშეებმა განაცხადეს, რომ ბევრი თამაში უჭირდათ".
მრავალფეროვანი უნარები: ყველა ამ თამაშის გადაჭრა მოითხოვს სხვადასხვა უნარებსა და სტრატეგიებს.
ექსპერიმენტები: თამაშებს გააჩნია სურვილისამებრ ექსპერიმენტის რეჟიმი, რომელიც საშუალებას აძლევს ადამიანებს ითამაშონ მათთან ერთად ისე ინტენსიური „ნაბიჯ ლიმიტის“ გარეშე, რაც მათ შეუძლიათ.
დამამშვიდებლად რთული: თამაშები საკმარისად რთულია, რომ დღევანდელი სასაზღვრო მოდელების დამარცხება შეუძლებელია.
How well do AI systems do? The benchmark is split into seven tiers with tier 1 being the easiest and tier 7 the hardest. 21 games have been released publicly with the remaining held back. Most of the games have multiple levels and the number of available actions for players to take at each step ranges from 2 all the way up to 34.
Opus 5 and Fable 5 with Claude Code are the best overall models, followed by GPT-5.5
Only Opus 5 and Fable 5 were able to beat any tasks (0.2) in (Tier 7). Opus 5, GPT-5.5, and Kimi K3 were able to beat some tasks in Tier 6 when given access to a harness (e.g, Claude Code).
GLM-5.2 and Gemini 3.1 Pro were able to beat some levels in Tier 4.
Overall, this seems really hard!
რატომ არის ეს მნიშვნელოვანი - კრეატიულობისა და აღმოჩენის მარიონეტები: მსგავსი ტესტები არის კრეატიულობის წინაპირობის იზოლირების მცდელობა, რომელიც შეძლებს ავტონომიურად აღმოაჩინოს სასარგებლო დაუსაბუთებელი რამ ახალი სიტუაციების შესახებ, რომლებშიც აღმოჩნდებით. როგორც ეს ტესტი აჩვენებს, ზოგიერთ სასაზღვრო მოდელს უკვე შეუძლია შეასრულოს საკმაოდ შთამბეჭდავი საქციელი ადამიანებთან შედარებით 7 საკმაოდ ცუდია იმ ფაქტთან შედარებით, რომ ცალკეულმა ადამიანებმა შეძლეს ტესტების 100% მიღება). ჩემი ვარაუდით, ჩვენ მივაღწევთ ადამიანურ პარიტეტს DiG-ს სკამზე 2027 წლის შუა რიცხვებისთვის, რა დროსაც უნდა ველოდოთ ისეთი რამ, როგორიცაა რეკურსიული თვითგაუმჯობესება, სერიოზულად დაიწყება.
წაიკითხეთ მეტი: DiG-bench: Discovery in Games (GitHub, PDF).
ითამაშეთ თამაშები და ნახეთ ლიდერბორდი ოფიციალურ საიტზე (digbench.ai).
***
მიიღეთ რეკურსიული თვითგაუმჯობესების შეგრძნება ბრაუზერზე დაფუძნებული ამ თამაშის თამაშით:
…ქუქიების დაწკაპუნება, მაგრამ სინგულარობისთვის…
აქ არის სახალისო თამაში ხალხისგან Paradigm Research, რომელიც მიზნად ისახავს სიმულაციას, თუ როგორია აწარმოოს კომპანია, რომელიც აშენებს AI სისტემებს, რომლებსაც შეუძლიათ რეკურსიული თვითგაუმჯობესების უნარი. თუ თქვენ თამაშობთ თამაშს, შეგიძლიათ მიიღოთ კარგი შეგრძნება იმის შესახებ, თუ როგორ ურთიერთქმედებენ ხელოვნური ინტელექტის კვლევის სხვადასხვა კომპონენტები, დაწყებული, როგორ დააბალანსებთ მკვლევარებში ინვესტიციებს და გამოთვლებს, როგორ და როდის უნდა ლიცენზირებული მონაცემები და სხვა. გაფრთხილდით, ძნელია - მაგრამ ისევ ასეა სასაზღვრო ხელოვნური ინტელექტის განვითარება.
რატომ აქვს ამას მნიშვნელობა: უკეთესი ინტუიციის შემუშავება რეკურსიული თვითგაუმჯობესების შესახებ არის ეგზისტენციალური მნიშვნელობა ჩვენთვის ყველასთვის; მსგავსი თამაშები გაგვიადვილებს მსჯელობას ამ ტექნოლოგიისა და მის მშენებელ ლაბორატორიებზე.
ითამაშეთ თამაში აქ: RSI Simulator (Paradigm Research).
***
AI systems are showing early signs of scientific research taste:
…Inherent post-trains an open weight model into an AI scientist that supervises a frontier model…
Taste is a hard thing to quantify but an intuitive thing to sense, as any of us know who have sat in a well-designed room, looked at someone wearing a particularly good fit, or read a research paper that asks just the right questions. Now, researchers with AI startup Inherent have published a paper showing how they are building Faraday, an AI scientist model that they hope can develop some taste in terms of research.
What they did: The company built a supervisory harness and relatively small LLM which sits on top of large, proprietary frontier models, and controls them in a way that improves their effectiveness at science. (In some ways, this is a capabilities-centric version of the scalable oversight problem).
To help them train and evaluate the system they assemble a dataset (”Replica”) consisting of research papers that have key graphs or results missing from them, then they see how well AI systems can autonomously do experiments that fill in the blanks, and they continuously train a small supervisory model (”Faraday”) via GRPO on well-designed fill-ins to achieve better and better results.
Faraday is a 27B model that uses a coding agent (OpenAI Codex) as an underlying tool and is post-trained on top of Qwen-3.6-27B.
What Replica consists of: Replica is a set of 100 ML and AI-for-science papers published between 1990 and 2026. The authors convert this dataset into a set of 310 replication tasks by knocking out individual results. “For each task, we use Claude Opus 4.7 prompted with a meta-rubric to generate a task-specific grading rubric,” they write. They then use a Codex-based Judge model to provide “an overall reward and per-turn credit assignment weights, which are used to train the Faraday agent using a modified version of GRPO.”
Results: Faraday using Codex is able to beat standard Opus 4.8 and GPT-5.5 on some replication tasks, exceeding their performance “on 73% of in-distribution ML tasks, and on 60% of held-out AI-for-science tasks, according to our rubric-based judge.”
“We achieve a comprehensive uplift in performance compared to the base Qwen model, on both train and test tasks,” they write.
Why this matters - the better systems like Faraday get, the higher the chance AI systems will become capable of recursive self-improvement: These days, most high-signal AI evaluations are trying to capture some property of creativity and intuition and Faraday/Replica is the same. The better AI systems get at this, the more likelihood we can assign to the idea that AI systems will imminently become capable of building themselves.
“The skills that allow Faraday to fill in vaguely-specified details may be the very same skills that would allow it to advance the state of the art by designing its own experiment,” the company writes. “The skills Faraday acquires – deciding what to investigate, scoping experiments to a budget, and judging a replication – compound with advances in frontier coding models. One might hope that a single post-trained outer agent can track the frontier as better models are released, at least over some time period.”
Read more: Training AI Scientists to Replicate Research (arXiv).
***
Mark Zuckerberg seems to be a technological pessimist:
…Zuck’s big essay on AI seems to ignore or elide or not confront what AI systems capable of invention mean…
Mark Zuckerberg has written an essay called “The Future is for Everyone“ that serves as something of a manifesto for how he and Meta are approaching the development of AI systems. The core idea inherent to Zuck’s strategy is to massively proliferate AI capabilities to everyone on the planet in a bid to avoid concentrating power and creating tyranny in a small number of players. It’s a broadly sensible idea except for the fact that superintelligences capable of inventing new ideas might want to do different things to what Mark Zuckerberg proposes and on this crucial area his essay is silent.
Mark’s view: “The defining questions of our age are who will have access to superintelligence and what will we direct it towards,” Zuckerberg writes. “We propose a philosophy based on individual empowerment as the source of prosperity, invention as the primary purpose of superintelligence, and balance of power as the foundation of safety.”
Meta’s goals and beliefs:
“Everyone will have an exceptionally capable personal agent that understands you, your goals, and everything you care about.”
“Everyone will have incredible tools for creation to express your ideas.”
„ყველას ექნება ძლიერი ინსტრუმენტები ახალი ბიზნესის შესაქმნელად და ეკონომიკა უფრო სამეწარმეო გახდება.
„ყველას ეყოლება პერსონალიზებული დამრიგებელი და მწვრთნელი, დოქტორის ხარისხით ყველა საგანში და შეუზღუდავი მოთმინებით, რომელიც დაგეხმარებათ გაიგოთ ყველაფერი, რაც გსურთ.
ყველა ისარგებლებს მეცნიერული მიღწევებით და შეძლებს თავისი წვლილი შეიტანოს სამეცნიერო პროგრესში.
”ყველას ექნება უფასო ან ხელმისაწვდომი წვდომა ამ ინსტრუმენტებზე.”
გამოტოვებული კითხვა: ამ სტატიის ის ნაწილი, რომელიც ყველაზე ნაკლებად მესმის, არის ცუკერბერგის მიერ ხელოვნური ინტელექტის სისტემების ერთობლიობა, რომელსაც შეუძლია გამოიგონოს ინდივიდუალური გაძლიერება. თხზულება სავსეა ისეთი საგნებით, რომლებიც, როგორც ჩანს, ვარაუდობენ, რომ ეს ყველაფერი შეფუთულია, მაგალითად:
„მიუხედავად იმისა, რომ კითხვების რაოდენობა, რომელსაც ადამიანს შეუძლია დაუსვას დღეში შეზღუდულია, ძვირფასი ნივთების რაოდენობა, რომელსაც სუპერინტელექტი შეუძლია გამოიგონოს თქვენი მიზნების მისაღწევად, შეუზღუდავია“.
„რაც უფრო მეტი სუპერინტელექტი ემსახურება როგორც გამოგონების ინსტრუმენტს, მით უფრო სავარაუდოა, რომ ინდივიდუალური შესაძლებლობები აჭარბებს ავტომატიზაციას და მომავალი უკეთესი იქნება ადამიანებისთვის“.
"მალე ყველას ექნება გამოგონების ზესახელმწიფოები."
"რომელ შედეგს მივიღებთ, დამოკიდებულია ბალანსზე მიმდინარე ბალანსზე ავტომატიზაციასა და მეორე მხრივ ინდივიდუალურ გაძლიერებასა და გამოგონებას შორის."
რატომ აქვს ამას მნიშვნელობა - ამ ყველაფერში დაკარგული კითხვაა: „იმუშავებს თუ არა სისტემა, რომელსაც შეუძლია ზეადამიანური გამოგონება, მხოლოდ იმ ადამიანების ინდივიდუალური გაძლიერების მიზნით, რომლებიც მასზე ნაკლებად შეძლებენ გამოგონებას?“. რა თქმა უნდა, ეს არის მთავარი კითხვა? მე არ ვარაუდობ, რომ ზეადამიანური გამოგონება უზრუნველყოფს რაიმე სახის ავთვისებიან არსებას, რომელიც დამოუკიდებელია ადამიანებისგან. პირიქით, მე ვარაუდობ, რომ ძნელია ზეადამიანური გამოგონების უნარის მქონე სისტემის შეჯერება ისეთ რამესთან, რაც ფუნდამენტურად არ ცვლის ძალთა ბალანსს მსოფლიოში ისეთი გზებით, რომლებიც დამაბნეველი და ძნელი დასაბუთებულია. როგორც ჩანს, ცუკერბერგი ასკვნის, რომ ამ სისტემების გავრცელება გამოიწვევს ძალაუფლების საწინააღმდეგო მყიფე ბალანსს სუპერდაზვერვით აღჭურვილ ადამიანებსა და კორპორაციებს შორის. ეს, რა თქმა უნდა, ერთ-ერთი პოტენციური შედეგია, მაგრამ მე ვცდილობ დავინახო, როგორ არის ეს წინასწარი შედეგი.
დაწვრილებით: მომავალი ყველასთვისაა (მეტა).
***
ტექნიკური ზღაპრები:
პირველი არკოლოგია
The Arcology was built for machine-subjective millennia, but to the humans its construction spanned a year. It was so vast and so complicated that watching it grow was akin to seeing plants rise up from bare dirt in fast-forward; jerking and growing in fits and starts, each of which spanned kilometres. Parts of it came alive while additions were added; rumors say some of its first halls to light up were reserved solely for computers to coordinate the construction of its next phases. When machines broke down determinations were made as to how valuable they were; if precious they would be taken nearby for repairs and returned to the site, but if below some threshold they were killed and stripped for parts where they had broken, then used to build the structure.
At night, an eerie ringing came on the air near it, both the sound of wind moving through its spindly and yet-unbuilt edges, and also the fans and hum of its slow dreaming computation, and finally the sound of the machines working through the night moving so quickly that they cut and tore the air into unnatural screams. To lie awake and hear an alien sound that spoke of your own successors must have been a strange thing indeed for the humans that lived within earshot.
Things that inspired this story: The construction of the pyramids; Gaudi’s Sagrada Família; tombs and future tombs.
Science AI; RSI სიმულატორი; და ზუკის ტექნოლოგიური პესიმიზმი

კომენტარები
ჯერ არავის დაუწერია კომენტარი. იყავი პირველი!
კომენტარის დასატოვებლად გთხოვთ გაიაროთ ავტორიზაცია ან დარეგისტრირდეთ.