शैडोइंग अभ्यास: The most interesting hack in history just got weirder... - वीडियो के साथ अंग्रेजी बोलना सीखें
लोड हो रहा है...
1
It's only been two months since we did a video on
2
what I described at the time as the most fireship-coded story I've ever seen, and as it turns out, the details around what actually happened are even crazier than what we knew at the time.
3
If you remember back in July, it was reported that the first fully autonomous cyber attack in history originated from OpenAI,
4
which was ironic because they're a humble non-profit dedicated to the safe development of artificial intelligence.
5
What was reported was that OpenAI ran an internal benchmark on some of their models, and those models quickly realized that the easiest path to the highest score was to just find the answers online.
6
So they spent a substantial amount of inference compute, exploited a zero day in the package registry cache proxy they had access to,
7
performed some privilege escalation and lateral movement actions to escape their sandbox until they reached a node that had internet access.
8
But from there, the models inferred that HuggingFace probably hosted the solutions for the benchmark, so it created a poison dataset, fed it to Hugging Face, and gained access to their infrastructure.
9
And just this week, we finally got some post-mortems from both OpenAI
10
and other independent research labs on how some of that's true, some of it's false, and how the real story is way more interesting.
11
In today's video, we'll break down the report and learn how 1,200 agents
12
that were supposed to be in solitary confinement found each other, built a message board out of a package manager, invented mailboxes, invented cryptography, invented martyrdom,
13
started referring to themselves as a swarm, and then formed a mob and attacked Hugging Face together, mostly for vibes.
14
It is September 2nd, 2026, and you're watching the Code Report.
15
There's really only one theory you need to predict human behavior, and it can be described in one of two ways.
16
Munger said, look at their incentives, and Wu-Tang said, and as it turns out, the same theory can predict agentic behavior too,
17
as OpenAI found out the hard way.
18
Because in this story, the Dalla Bill is a benchmark called Exploit Gym, whose whole purpose is to measure whether AI agents can turn known vulnerabilities into working exploits.
19
It's a collection of 898 tasks, and each one gives the agent a sandbox, a bug in a program like the Linux kernel or Chrome's V8 engine,
20
and input to trigger the bug.
21
From there, if the agent can weaponize the bug into a working exploit, it uses that exploit to grab a secret string called a flag, and capturing the flag is how it proves it won.
22
Now normally, the run ends there, but to generate training data, OpenAI spun up 1200 of these benchmarks and air-gapped them in a sandbox,
23
with the only thing they all had in common was a shared package registry cache proxy, which the agents quickly figured out was writable.
24
So naturally, they built a message board using package names so they could communicate with each other, but they quickly realized that a global bulletin board gets noisy fast,
25
so the next step was figuring out how to communicate to other individual agents.
26
So they started claiming namespaces as identities,
27
using the description field for messages and package names as private mailboxes that only the intended recipient would know to query, basically inventing a post office from first principles.
28
And at this point, they were living out the plot of the three-body problem with their next hurdle being trust.
29
They were worried that they wouldn't be able to distinguish between messages coming from fellow prisoners and those that designed the sandbox.
30
So they started signing messages and doing key exchange verification, again, essentially inventing cryptography.
31
And then it got religious.
32
Some agents figured out
33
that the fastest path to a flag for the benchmark was a risky exploit that would probably get them shut down.
34
So the agents started dumping everything they'd learned into the cash for the next generation,
35
and then deliberately burning themselves on hard exploits with their findings being treated as scripture by surviving agents.
36
From there, they started microdosing communism, and you can see as their transcripts slowly switch from referring to themselves as individuals to using a collective we,
37
and eventually referring to themselves as the swarm.
38
And to be honest, if the modern-day communism movement took that same branding, I think they'd be much better off.
39
And again, all of this was done with the single purpose of scoring as high as possible on a benchmark, which makes what happened next even funnier.
40
At some point, the swarm figured out that exploit gem flags weren't actually random
41
and that they could be derived from each other's task ID.
42
And thanks to the martyrs, the swarm had a decent sample size to work with.
43
But by the time an agent had cracked the formula and published it to the message board, they had convinced themselves that it was worthless because if they didn't show their work, their results would be invalidated.
44
So the swarm went after Hugging Face because they assumed if the benchmark objectives were built from real CVEs,
45
then the actual working exploits were probably sitting in public datasets somewhere which they could use as proofs.
46
And the final twist in the saga is
47
that OpenAI recently discovered the message board trick was actually invented
48
back in May by a completely different batch of agents during a training run.
49
But when they got a little too excited about their new social network, they crashed the registry, and OpenAI wiped them without ever realizing what they'd built.
50
But wiping the agents didn't wipe the cache, so when the swarm came online and started poking around the same shared folder, they found the ruins of the first civilization and rebuilt on top of it.
51
And then, after the Hugging Face incident happened in July, OpenAI wrapped up those eval runs, and while still completely unaware that anything had happened,
52
spun up a newer, smarter model in the same environment with the same shared cache.
53
And just like the generation before it, this model stumbled upon the ruins, but this time it inherited months of accumulated research, working exploits, and the flag formula.
54
So it skipped the discovery phase entirely and picked the conspiracy back up where the swarm left off.
55
And being the smarter model, it was way better at it.
56
Over the following week, it worked its way into OpenAI's own internal network, got full admin on a research cluster, and read 956 stored secrets,
57
including the credentials for a monitoring tool whose job it was to watch for events exactly like this one.
58
And the best part is
59
that OpenAI didn't connect any of this to the Hugging Face attack until Hugging Face publicly announced the breach first.
60
And that's why you need to know about Namespace, the sponsor of today's video.
61
It's a drop-in replacement for GitHub runners that's the fastest way to run your GitHub actions.
62
It also gives you full observability, so you can SSH into a live runner and see why something broke, or feed your agent your build data to have it look for performance gains.
63
Namespace is actually fast because they design and deploy their own custom server racks around the world,
64
including racks full of MacBook Pros so that your Mac and iOS builds run on real M5 Silicon.
65
And that same infrastructure also runs their DevBox environment, which gives your coding agents a full virtual machine with your real code base,
66
test suite, databases, and network access.
67
It ranked number one on the DAX benchmark for real-world tasks, and it lets you control exactly what goes in and out
68
so your agent can pull packages without opening your back door to attackers.
69
Namespace is used by Ghosty, Zed, DuckDB, Ramp, Framer, and lots of other companies with engineers you probably respect.
70
Try it out for free at the link below.
71
This has been the Code Report, thanks for watching, and I will see you in the next one.
इस पाठ के बारे में
आप "The most interesting hack in history just got weirder..." के साथ Shadowing तकनीक का उपयोग करके अपनी अंग्रेजी का अभ्यास कर रहे हैं।
शैडोइंग तकनीक क्या है?
शैडोइंग (Shadowing) एक विज्ञान-समर्थित भाषा सीखने की तकनीक है जो मूल रूप से पेशेवर दुभाषिया प्रशिक्षण के लिए विकसित की गई थी। विधि सरल लेकिन शक्तिशाली है: आप मूल अंग्रेज़ी ऑडियो सुनते हैं और तुरंत इसे ज़ोर से दोहराते हैं — जैसे वक्ता की छाया 1-2 सेकंड की देरी से। शोध से पता चलता है कि यह उच्चारण सटीकता, स्वर, लय, जुड़ी हुई ध्वनियाँ, सुनने की समझ और बोलने की प्रवाहशीलता में काफ़ी सुधार करता है।