शैडोइंग अभ्यास: We Let $200 GPT-Pro Think for 7 Days on Unsolved Math - वीडियो के साथ अंग्रेजी बोलना सीखें
लोड हो रहा है...
1
So as some of you will have seen, the rate at which AI is solving research level maths problems has accelerated drastically in the last few months.
2
Problem after problem after problem is supposedly being solved by these large AI teams like OpenAI and Anthropic.
3
You know, recently we saw that a larger proportion of the zeros of the Riem
4
and Zeta function have been proven by Claude to lie on the critical line, which is a development of a long line of work, one of the main ones being Brian Connery's work,
5
who showed that an initial proportion of the zeros of the Riem and Zeta function lie on this critical line.
6
And there's a really funny quote from a talk that I saw once.
7
In 1989, I proved that at least two-fifths of the zeros are on the line.
8
When they offered this million-dollar prize, I wrote to them and said, well, I've already proven two-fifths of this, so I think you should give me $400,000.
9
But the thing is, the majority of these really difficult problems have been solved with harnesses
10
or models which haven't been released to the public yet.
11
Nonetheless, in the last few videos, we've seen some serious improvements in the performance of these models, solving PhD kinds of problems that I've come across in my research,
12
and today we're going to be taking it, like, another step further.
13
Because, like I mentioned last time, a member from OpenAI's codex team reached out to me
14
and basically planned a meeting where he explained to me how
15
I might be able to tackle some of these problems using codex in a more kind of refined and systematic way.
16
Since then, I've also had a member of Anthropic reach out, so I've actually now got both of the best models, like the $200 subscriptions, which is absolutely awesome, so thank you so much to both of you guys,
17
and they've even said to me, you don't have to be biased, so honestly that's just really awesome.
18
But in any case, today we're going to be using this in order to have $200
19
chat GPT think for well over a week on a class of open mathematical problems in random matrix theory
20
that would be really cool to see an answer to, even if it's an answer that's written in clanker language, okay?
21
Because one of the problems with having such powerful models is the more powerful they become, the harder it is to kind of interpret in a human way what exactly it is that they're doing.
22
We saw that last time
23
when we gave Anthropic $100 of AI tokens to see whether it could reproduce a research file that has any interesting results.
24
And honestly, it felt a bit like I'd given Dario $100 and then the client call was like, $100 please.
25
It took the money, said a bunch of gibberish.
26
It was actually kind of true.
27
So it's not that it did things wrong, but maybe it didn't actually give the types of results that I thought were interesting.
28
But in the last week, I've had codecs running nonstop.
29
There's been lines of reasoning that have gone on for one day, lines of reasoning that have gone on for two days.
30
We're well over seven days now of the clanker thinking in
31
understanding something that's quite important with regards to the Riemann zeta function, but that's actually approached from a completely different angle using tools of random matrix theory.
32
So I'm going to be discussing basically how this system works, and we're going to be looking like kind of inside the mind of the clanker that's been thinking for ages, because it does feel a bit like it's been in some long trance.
33
I was reading the reasoning, but generally speaking, it was quite hard actually to understand what exactly it was doing.
34
But anyway, we're going to load up the project now, and I'm going to show you exactly how it's been done.
35
Let's go.
36
So as you can see, we're using the Mac application for codecs, and actually what it's been doing is it's been working in this kind of subfolder, but as you can see here, we've been doing a whole bunch of research,
37
basically starting from this $100 Clawed tokens that we gave it last time, and I've been continuing this essentially with a whole bunch of MD files.
38
The Clankers are trying loads of different things, I mean, as you can see, it's doing loads of like tracking of the things it's tried, where it's gone.
39
I'm basically just trying to collect with all of the tokens that I've been given access to through both companies, all of the different paths.
40
So to try and search to see if there is a way of doing it.
41
This is some study of the modes of the distribution that I'm trying to understand.
42
At this point, I'm just trying to see whether the clanker can agree
43
that it understands something and then to try and systematically read that path that it does, as opposed to reading the thousands of things that it's done.
44
I mean, just look at all of these files.
45
It's just ridiculous.
46
There's just no way I can read this at this point.
47
So I'm taking the method a bit to the extreme compared to what we've done in the last few years, because, you know, we've seen these models have become much, much better at doing these types of tasks,
48
have it go off in all of the directions.
49
And then when it finds a direction that's useful for me to then follow the reasoning of that direction, to have the models verify their reasoning before.
50
So this is the chat after Codex has crashed.
51
It was thinking for about 24 hours and it had edited loads of files
52
and added them to the git folder
53
so all of its chain of thought was basically in editing these .md files
54
and I asked it to essentially summarize what it had done.
55
So there it took three minutes.
56
So this number, just looking at it, is quite interesting because I think what it's basically doing is positioning the mode
57
and it's giving you an actual location for it.
58
As you can see, we kind of had it think for way longer
59
and it was giving all kinds of explanations of what exactly was going on.
60
So the problem that I was basically asking it to look at, which is to do with why a certain distribution, as you go up in matrix dimension, there's like a transition where it goes from having one mode to having two bumps.
61
And the thing that most people are interested in is whether it still has two bumps when the dimension goes to infinity.
62
After reading this for quite a while, so there we've got another hour, there we've got 17 minutes, which is nothing, then we've got one minute, there we've got 13 hours of it basically working on this problem,
63
it started to feel like what it was basically giving was like some sort of machine proof.
64
So it was essentially incomprehensible to me.
65
So it gave all of these summaries, so this is still its reasoning, still its reasoning.
66
It was literally just thinking for ages.
67
This is the 13 hours of chain of thought.
68
So I'm just going to scroll down.
69
It just did loads of stuff, okay?
70
This is what I mean.
71
It's kind of just, I don't know whether it's just hallucinating.
72
I don't think, I think there is a lot of truth in what it's doing, but to me, I can't, it's just not possible, okay, for a human to read through all of this stuff.
73
Then I asked it for a recap after this 13 hours of thinking, which was a lot.
74
Basically it was just giving me loads of numerical numbers, like, and then it basically says that they're not actually numerical, but they're kind of like from inequalities,
75
basically saying that it was a form of computer assisted proof.
76
Yeah, so current work is largely rigorous computer assisted analysis, not merely numerical experimentation now.
77
Essentially it was taking polynomials and trying to understand, write them in a certain basis, and then showing that the coefficients like have a certain sign,
78
and then that can kind of deduce information about the question that we're asking to do with these modes.
79
That's my understanding of it, okay?
80
But nonetheless, it kind of summarized something quite interesting, which is this.
81
So this is a general n expression of something that you need to evaluate.
82
It did manage to come up with a candidate justification as to why n equals 9 was special.
83
But then it couldn't really prove rigorously that that was actually the case.
84
Kind of said that I was not really that interested in all of these kind of machine-assisted proofs, I really just wanted like an analytic description of what was different at n equals 9 compared to the other ones.
85
Here I was just trying to set up a plan, so this is a bit like where you can essentially just ask Codex to set up a plans.md file
86
or some sort of like systematic way, this was guidance that was given from the guy from Codex, but you have a discussion with the model in order to get this done, so here you see it had like a research tracker
87
and then it had this like analytic program
88
which is where it actually has a bunch of tasks that it's going to try and solve.
89
So then it basically outlined which things are possible, so medium to low probability that it could prove that there's two maxima in the limit, which is kind of one of the most interesting problems.
90
So then it summarized all of the known results, and this is where it started to work on thinking for a very very long time, and it's literally been thinking for multiple days now,
91
like, and we're going to see what it's done.
92
So this is the document that the Clankers have produced, which is summarizing this large hallucination, which is probably actually legit,
93
I am only joking, but it is funny seeing so much stuff be generated.
94
So we're gonna go through it.
95
It's made a 16-page document with a final assessment.
96
All hell breaks loose when you just let the clankers do it, but maybe that's how you should do it.
97
We don't know yet.
98
So it's worth trying.
99
But okay, so it's basically just outlining the research tracker.
100
So you can see it's got 321 markdown research notes.
101
It's just made loads of stuff.
102
Two different logarithmic maxima.
103
The density therefore is generated by the switching set for at least two.. oh my goodness this doesn't make any sense.
104
The one constraint formula.
105
Okay, so this is cool.
106
In principle, if I understand it right, which I'm not even sure if I do, some number of constraints when you reduce, when you try and get this radial density, and the problem is
107
that what it ends up counting is like intersection points of like curves onside like an n-dimensional torus or something, sum something over those points, which is essentially just a constraint,
108
a set of constraints, and it's saying that in the large end limit you only get one constraint, which would be cool.
109
And then here I guess it's trying to probably show why you get two peaks.
110
Maybe this is the minimum, these are the two maxima.
111
Of course, two strict local maximas separated by local minimum.
112
So I guess that's just the diagram makes that kind of obvious.
113
So it's basically saying that you can get the red curve as an integral, and the thing that you're integrating, you can write it explicitly, supposedly.
114
I'm surprised that it's got a bump here.
115
It's saying that it's doing the marginal of the JPDF, essentially like numerically computing the thing that you've proven analytically.
116
It's not really true, I feel.
117
I shouldn't have a bump there.
118
Again, it cannot explain the new low-order radio station repair.
119
These results move to causal questions away from the local singularities and towards global connection coefficients.
120
I mean, I am talking to an alien right now, I have to say.
121
I mean, I do feel like I'm talking to someone who's done a lot of stuff.
122
I feel like if the proof looked like this, I mean, I don't know if you guys saw Terry Tao made a talk.
123
And I thought it was a really good talk, actually, because one of the, like, it's not just about proving things.
124
It does need to be comprehensible.
125
I mean, is it that useful if we get approved of stuff and it's in Clanker language?
126
Like, who cares?
127
I mean, it's not really that interesting in a way
128
because it hasn't actually given you a more interesting underlying structure that shows that that's why it is the way it is.
129
But we didn't stop there because, you know, I actually wanted to see even more of what the Clanker was able to do.
130
So we gave it an even more difficult problem.
131
Not that it solved the previous one.
132
And the main goal of this one was to prove whether this two-hump phenomenon stays
133
when you take the dimension of the matrix to infinity.
134
So as you can see here, it's been thinking for four days.
135
Which was a long time.
136
I had my computer running 24-7 with a piece of software called...
137
What was it called again?
138
It's called Amphetamine.
139
So it's actually just like a piece of software
140
that generalizes some sort of function that you can put in the terminal that like lets your Mac go on forever.
141
It's just got slightly better settings to that Mac command.
142
But yeah, my computer's been on for a long time, you know, overnight with the fans on.
143
And what's come out is quite funny.
144
Now, I'm not trying to take the piss, okay?
145
So maybe the- the Clankers are probably cleverer than I am, for sure.
146
Okay, well I'm not just gonna roll over and like take it like that, but in a lot of ways they are, okay?
147
In other ways they're not.
148
And you could argue this is one.
149
So, after thinking, so it also thought for a long time before this, I mean I'm not gonna like, summarize every single thing that it did, but pretty much it thought for multiple sessions of like,
150
loads of hours, and after two days it said, Yes, I've done it.
151
I've proven in a rigorous sense that there's at least two distinct modes.
152
And its proof is like these inequalities.
153
Do you see what I mean?
154
I mean, are you for real?
155
The completed outward certificate proved that T 1.5 0.4, so these are, I guess,
156
some like, locations that you're evaluating something is bigger than 0.00012659100.
157
I'm literally talking to a clanker right now.
158
Here we've got some more after it thought for a bit longer.
159
It gave me some other certified analytic inequalities.
160
I mean maybe people do care.
161
I mean let me know if you're interested in this.
162
I don't really think it's that interesting if you're just you've got two inequalities.
163
To me it feels like numerical evidence to be honest.
164
Even though I suppose it must be fundamentally different if you're somehow rigorously showing something, it's just like taking brute force to the next level.
165
But with this, I had this going on the side, like while I was writing my thesis, I had codex going on this thing, just to see whether, if you just let it go and think forever, whether it gets anything.
166
And it thought for multiple times, so this was one session where it thought for four days, I had some other ones where it thought for like three days, and I had one at the beginning that was like one or two days that it thought for,
167
so I've been recording this for a long time.
168
But yeah, that was what we ended up with.
169
We ended up with this.
170
So overall it's definitely been interesting seeing how Codex can actually
171
work on a project systematically with all of the different research tasks as goals
172
and see how it can generate all of these ideas as MD files
173
and all of its reasoning and subsequent work so that it can keep this context for future research.
174
Of course now
175
that I've been given access to both of these models we're going to be putting Opus
176
and Fable in on this project as well
177
and see how far they can go in trying to understand all of these problems
178
because it's a long list of problems all about one thing basically in random matrix theory
179
And it'll be cool to see if the Clankers can slowly but surely figure it out.
180
That would be very interesting.
181
I want to say thank you guys so much for all of the support, like all the channel members and anyone who watches this far.
182
But as always, I hope you enjoyed the video and have a good one.
इस पाठ के बारे में
आप "We Let $200 GPT-Pro Think for 7 Days on Unsolved Math" के साथ Shadowing तकनीक का उपयोग करके अपनी अंग्रेजी का अभ्यास कर रहे हैं।
शैडोइंग तकनीक क्या है?
शैडोइंग (Shadowing) एक विज्ञान-समर्थित भाषा सीखने की तकनीक है जो मूल रूप से पेशेवर दुभाषिया प्रशिक्षण के लिए विकसित की गई थी। विधि सरल लेकिन शक्तिशाली है: आप मूल अंग्रेज़ी ऑडियो सुनते हैं और तुरंत इसे ज़ोर से दोहराते हैं — जैसे वक्ता की छाया 1-2 सेकंड की देरी से। शोध से पता चलता है कि यह उच्चारण सटीकता, स्वर, लय, जुड़ी हुई ध्वनियाँ, सुनने की समझ और बोलने की प्रवाहशीलता में काफ़ी सुधार करता है।