1
00:00:09,750 --> 00:00:11,830
Welcome to the Road to Accountable AI.

2
00:00:12,030 --> 00:00:15,380
I'm Kevin Warbeck, Professor of Legal
Studies and Business Ethics at the Wharton

3
00:00:15,380 --> 00:00:17,120
School of the University of Pennsylvania.

4
00:00:17,770 --> 00:00:21,509
For decades, I've studied emerging
technologies from broadband to blockchain.

5
00:00:22,255 --> 00:00:27,034
Today, artificial intelligence promises
to transform our world, but AI won't reach

6
00:00:27,034 --> 00:00:31,425
its potential without accountability,
mechanisms to ensure it's deployed in

7
00:00:31,425 --> 00:00:33,715
responsible, safe, and trustworthy ways.

8
00:00:34,235 --> 00:00:38,185
On this podcast, I speak with the experts
leading the charge for accountable AI.

9
00:00:40,805 --> 00:00:44,774
Whether to restrict open weight foundation
models is a big topic of discussion

10
00:00:44,874 --> 00:00:47,165
in AI safety and policy circles.

11
00:00:47,165 --> 00:00:51,264
Thanks But not one that the general
public and business community are

12
00:00:51,264 --> 00:00:54,444
necessarily focused on my guest,
Kevin Bankston from the center for

13
00:00:54,445 --> 00:00:59,164
democracy and technology talks about
why this issue is so important,

14
00:00:59,494 --> 00:01:03,944
both to the future of innovation in
AI and to our ability to address.

15
00:01:04,045 --> 00:01:09,015
Serious AI risks, as you'll hear, Kevin
has strong views and in full disclosure,

16
00:01:09,045 --> 00:01:13,384
I signed on to an open letter that
CDT wrote on these issues, but in our

17
00:01:13,384 --> 00:01:17,504
conversation, we cover the arguments
on both sides of the issue and why

18
00:01:17,505 --> 00:01:21,524
this topic is actually so important
to understanding what the future

19
00:01:21,524 --> 00:01:23,645
path holds for the development of AI.

20
00:01:23,945 --> 00:01:25,845
Kevin, pleasure to have
you on the podcast.

21
00:01:26,424 --> 00:01:31,530
Before we get into talking about open
weight models, just a quick Tell us

22
00:01:31,530 --> 00:01:35,420
a little bit about your background
and for those who aren't familiar

23
00:01:35,420 --> 00:01:38,940
what CDT is, what's the work that
the organization is doing on AI?

24
00:01:39,800 --> 00:01:40,270
Sure.

25
00:01:40,309 --> 00:01:40,919
I'll start with CDT.

26
00:01:40,920 --> 00:01:46,849
CDT is, I would say, the premier
internet and technology policy think

27
00:01:46,849 --> 00:01:48,689
tank and advocacy organization in DC.

28
00:01:48,710 --> 00:01:52,580
We've been around for, I want to
say, almost 30 years at this point.

29
00:01:52,960 --> 00:01:57,890
Non partisan, non profit,
focused on protecting civil

30
00:01:57,890 --> 00:01:59,549
rights and civil liberties.

31
00:01:59,895 --> 00:02:03,395
In the digital world that
encompasses a whole lot of things.

32
00:02:03,404 --> 00:02:08,284
It's a broad spectrum organization, but
obviously right now AI is at the forefront

33
00:02:08,285 --> 00:02:10,365
of a lot of technology policy discussions.

34
00:02:10,764 --> 00:02:15,124
And I, um, with my colleague,
Miranda Bogan co founded

35
00:02:15,124 --> 00:02:16,795
the new AI governance lab.

36
00:02:17,190 --> 00:02:23,140
At CDT, um, which is focused on helping
develop best practices and standards

37
00:02:23,160 --> 00:02:29,829
around evaluating and mitigating AI risks
to help build a technically informed basis

38
00:02:29,840 --> 00:02:37,760
for eventual legislation and regulation,
taking advantage of not only my 20 odd

39
00:02:37,770 --> 00:02:41,130
years in civil society at organizations.

40
00:02:42,060 --> 00:02:43,679
Like and including CDT.

41
00:02:44,380 --> 00:02:51,140
After that 20 odd years, I was approached
by Facebook, now Meta, to basically

42
00:02:51,140 --> 00:02:55,000
help them figure out their path
forward on responsible AI development,

43
00:02:55,429 --> 00:02:57,589
which was an intriguing possibility.

44
00:02:57,750 --> 00:03:02,739
And so I spent about four years
building their AI policy team, being

45
00:03:02,739 --> 00:03:06,829
one of the founding senior leaders
of their responsible AI product team,

46
00:03:07,030 --> 00:03:10,780
which consults with all the other
product teams on their AI products.

47
00:03:10,820 --> 00:03:15,775
Thank you And did a lot of good work, but
obviously working in a big company like

48
00:03:15,775 --> 00:03:19,704
that in a lot of ways is very stressful
and challenging and time consuming.

49
00:03:20,075 --> 00:03:24,744
And I am an, at this point, a late
middle aged fella who just, more

50
00:03:24,744 --> 00:03:26,734
accurately, my wife just had a baby.

51
00:03:27,354 --> 00:03:30,054
And I decided for a variety of
reasons, but especially including

52
00:03:30,054 --> 00:03:32,174
that one, I was ready to leave.

53
00:03:32,334 --> 00:03:36,765
The corporate world and come back to civil
society and apply what I had learned.

54
00:03:36,955 --> 00:03:40,755
And also what Miranda, as my colleague
on my team at Metta had learned

55
00:03:41,075 --> 00:03:42,634
around responsible AI development.

56
00:03:42,644 --> 00:03:49,624
So now Miranda who did trail
trailblazing groundbreaking work at

57
00:03:49,625 --> 00:03:54,105
Metta in terms of addressing issues of
race and gender bias in ads delivery.

58
00:03:54,525 --> 00:03:57,375
She is the director of the
new governance project at CDT.

59
00:03:57,605 --> 00:04:00,575
While I'm a senior advisor to that
project and to the leadership at

60
00:04:00,575 --> 00:04:02,285
CDT on other AI policy issues.

61
00:04:03,040 --> 00:04:09,210
I currently teach the First Amendment
and copyright in regard to AI technology

62
00:04:09,210 --> 00:04:11,020
as an adjunct professor at Georgetown.

63
00:04:11,420 --> 00:04:15,470
Those are issues that CDT
has not spoken greatly about.

64
00:04:15,510 --> 00:04:19,620
We do address the First Amendment in
regard to this issue a little bit in the

65
00:04:19,620 --> 00:04:20,839
comments we're going to be talking about.

66
00:04:21,130 --> 00:04:24,030
But speaking generally, if I talk about
the First Amendment here, attribute

67
00:04:24,039 --> 00:04:28,750
that to me as Kevin the scholar,
not necessarily CDT's positions.

68
00:04:29,710 --> 00:04:30,020
Great.

69
00:04:30,020 --> 00:04:33,640
So this is why you are a perfect
person to talk about these issues.

70
00:04:33,660 --> 00:04:37,470
I want to ask you in particular
about the debate that's happening now

71
00:04:37,470 --> 00:04:39,600
around open weight foundation models.

72
00:04:39,600 --> 00:04:43,380
So first, for those who aren't necessarily
familiar, what are we talking about here?

73
00:04:43,949 --> 00:04:47,315
So this will require a little bit
of a history lesson about What

74
00:04:47,345 --> 00:04:52,105
open source software is before
we get to what open source AI is.

75
00:04:52,445 --> 00:04:58,845
So in the nineties, you began to see
the evolution of this new form of.

76
00:04:59,564 --> 00:05:05,064
Software development and licensing
around this open source concept where.

77
00:05:05,550 --> 00:05:11,830
You would have companies or even voluntary
projects with a bunch of contributors

78
00:05:12,130 --> 00:05:17,490
developing their code openly, as in
collaboratively, or when publishing their

79
00:05:17,500 --> 00:05:23,729
programs, also publishing the source code
itself so that other people could use that

80
00:05:23,749 --> 00:05:29,489
to build their own software, to modify it
for their own uses, et cetera, et cetera.

81
00:05:29,499 --> 00:05:32,839
Usually this is accomplished through
something called an open source license,

82
00:05:33,139 --> 00:05:38,000
which basically says, You can use this
and there are no use restrictions on it.

83
00:05:38,010 --> 00:05:38,750
We are essentially.

84
00:05:39,195 --> 00:05:41,845
Although we still own the copyright
in this technically, we are

85
00:05:41,845 --> 00:05:44,914
licensing it broadly for public use.

86
00:05:45,195 --> 00:05:49,594
So, why is this important
and, speaking generally, good?

87
00:05:49,635 --> 00:05:53,094
Certainly the internet and digital
technologies built around and

88
00:05:53,105 --> 00:05:57,995
on the internet rely an enormous
amount on open source software.

89
00:05:58,005 --> 00:06:02,974
At this point, 96 percent of all
codebases include some open source code.

90
00:06:03,445 --> 00:06:07,595
It's the original open sourcing of
the Netscape browser in the 90s, which

91
00:06:07,595 --> 00:06:12,234
led to the open source Mozilla Firefox
browser, which was a predecessor to

92
00:06:12,234 --> 00:06:16,374
later the Google Chrome open source
browser that enabled competition

93
00:06:16,595 --> 00:06:20,745
so that we didn't all just have to
use Microsoft Explorer on the web.

94
00:06:21,395 --> 00:06:26,095
Now, I want to be clear, open source
is not and will not be Always a

95
00:06:26,105 --> 00:06:27,875
silver bullet for competition.

96
00:06:28,185 --> 00:06:31,435
And it can be leveraged by
already powerful stakeholders

97
00:06:31,435 --> 00:06:32,375
to make them more powerful.

98
00:06:32,385 --> 00:06:36,505
Certainly it's enhanced Google structural
power, the fact that it controls

99
00:06:36,824 --> 00:06:41,824
ultimately the development of Android,
for example, but then again, that, that

100
00:06:41,824 --> 00:06:46,015
power has also helped counterbalance the
power of other people, other companies,

101
00:06:46,025 --> 00:06:50,234
like your Microsoft's and your Apple's,
and enabled a large number of hardware

102
00:06:50,234 --> 00:06:51,424
vendors to compete that wouldn't.

103
00:06:51,810 --> 00:06:55,880
Have had an operating system to
use and it's complicated ecosystem.

104
00:06:55,880 --> 00:06:59,310
But ultimately open source in the
context of regular software has

105
00:06:59,310 --> 00:07:00,750
been a boon for competitiveness.

106
00:07:00,950 --> 00:07:04,540
Also a boon for security in
the world of open source.

107
00:07:04,549 --> 00:07:07,860
They say enough eyeballs
make all bugs shallow.

108
00:07:08,239 --> 00:07:09,814
This is because if you have.

109
00:07:10,125 --> 00:07:13,075
Your source code interrogable by anyone.

110
00:07:13,215 --> 00:07:15,715
That means you have a lot of eyes
looking at it and finding flaws

111
00:07:15,715 --> 00:07:17,725
and helping you patch those flaws.

112
00:07:17,735 --> 00:07:21,994
That's why, for example, we have
the security world has generally

113
00:07:21,994 --> 00:07:26,495
been supportive of the open
sourcing and wide availability of.

114
00:07:27,075 --> 00:07:30,655
Penetration tools for breaking
into computer systems, not because

115
00:07:30,655 --> 00:07:33,085
they want to help people break into
systems, but because they want to

116
00:07:33,085 --> 00:07:37,505
help people detect how to break into
systems so they can fix those things.

117
00:07:37,905 --> 00:07:40,294
And so speaking generally, open
source has also been a boon.

118
00:07:40,615 --> 00:07:42,515
For our overall technical security.

119
00:07:44,044 --> 00:07:49,015
Now we have to get to AI though, and
open source is a little bit of a misnomer

120
00:07:49,025 --> 00:07:52,405
when talking about AI for reasons
I'll get to first, I'll have to back

121
00:07:52,415 --> 00:07:56,385
up and say, what are we talking about
when we talk about AI very generally.

122
00:07:56,720 --> 00:07:59,100
When we're talking about AI, we're
talking about machine learning,

123
00:07:59,360 --> 00:08:02,630
which is using software to look
at vast amounts of data to find

124
00:08:02,630 --> 00:08:04,360
patterns in it that are useful to us.

125
00:08:04,680 --> 00:08:09,000
In the current context, we are
usually talking about large language

126
00:08:09,030 --> 00:08:11,889
models or other generative models.

127
00:08:12,120 --> 00:08:14,899
And to break that down as simply as
possible and get to the answering

128
00:08:14,949 --> 00:08:16,510
the question of what are weights.

129
00:08:16,930 --> 00:08:22,470
I'll just say to make a large language
model, you take a whole lot of data.

130
00:08:22,935 --> 00:08:26,545
Often everything you can scrape
from the web, data you've

131
00:08:26,545 --> 00:08:29,955
licensed from other vendors, a
large amount of language, right?

132
00:08:30,215 --> 00:08:36,944
A large amount of language and feed it
through an algorithm that creates what

133
00:08:36,945 --> 00:08:42,454
is essentially a massive mathematical
space that sort of summarizes all that

134
00:08:42,455 --> 00:08:45,075
data, not directly, but rather each.

135
00:08:45,915 --> 00:08:49,135
Not even each word, each part of a
word, and they're called tokens in this

136
00:08:49,135 --> 00:08:55,964
context, gets its own little spot in
this sort of semantic map of all of that

137
00:08:55,964 --> 00:09:02,500
content, so that, for example, as it
has its ingested more and more data, it

138
00:09:02,500 --> 00:09:07,260
begins to not understand, but be able
to place close to each other related

139
00:09:07,260 --> 00:09:13,359
semantic concepts like dog and puppy, for
example, are close together in this map.

140
00:09:13,869 --> 00:09:19,410
What happens then is once you've trained
that model to create that mathematical

141
00:09:19,410 --> 00:09:25,275
map, which is what we call weights,
Because the, to put it in non technical

142
00:09:25,275 --> 00:09:29,074
language, like all of that semantic
content has been weighted to some extent

143
00:09:29,074 --> 00:09:30,935
or another in this mathematical model.

144
00:09:31,494 --> 00:09:35,684
You need the weights to be able
to infer anything for the model

145
00:09:35,685 --> 00:09:37,804
to actually give you an output.

146
00:09:38,905 --> 00:09:43,064
Open weights, and so now the
analogy to open source comes in.

147
00:09:43,584 --> 00:09:48,055
There, as with regular software,
there's both open and closed approaches.

148
00:09:48,725 --> 00:09:57,055
The closed approach is, for most of
its models, The chat, a good example

149
00:09:57,055 --> 00:10:00,545
is OpenAI and the way it treats most
of its most powerful models, which

150
00:10:00,545 --> 00:10:04,974
is it is available through them.

151
00:10:05,185 --> 00:10:11,005
You have to either go to their website
and pay the 20 bucks to access their

152
00:10:11,035 --> 00:10:17,405
user interface for chatting with
ChatGDP, or you're a developer who

153
00:10:17,405 --> 00:10:21,175
pays to access their API, their
application programming interface.

154
00:10:21,565 --> 00:10:26,285
To build apps or services, internal
or external on top of their model.

155
00:10:26,944 --> 00:10:31,825
Either way, though, they have
centralized control over the model.

156
00:10:32,395 --> 00:10:36,815
Users and developers do not have
access, direct access to the weights.

157
00:10:37,305 --> 00:10:38,795
And yeah, so that's closed AI.

158
00:10:39,565 --> 00:10:44,255
Open weights AI, and I hesitate to call
it open source for reasons I'll explain,

159
00:10:44,815 --> 00:10:52,055
is when you publish, at the very least,
the model software and the weights.

160
00:10:52,635 --> 00:10:58,375
So that another developer could deploy
that on their own infrastructure

161
00:10:59,155 --> 00:11:01,425
with their own modifications.

162
00:11:01,995 --> 00:11:06,115
So in many ways it is similar
to open source software because

163
00:11:06,125 --> 00:11:09,525
these open weights models are
typically offered under a license.

164
00:11:09,985 --> 00:11:15,515
That either that, that allow for free
redistribution and or don't have use

165
00:11:15,515 --> 00:11:20,345
restrictions and or otherwise give
you freedom to play with the model

166
00:11:20,345 --> 00:11:24,805
and build derivative models in a way
that you can't do with a closed model.

167
00:11:25,495 --> 00:11:26,374
I don't call that.

168
00:11:26,580 --> 00:11:28,630
Open source for a few reasons.

169
00:11:29,210 --> 00:11:35,780
Um, one, most of the components of the
model are not in the form of source code.

170
00:11:36,120 --> 00:11:40,240
And in fact, this critical heart of
the model, the weights, is not really

171
00:11:40,269 --> 00:11:41,830
interpretable by a human at all.

172
00:11:42,400 --> 00:11:46,050
Um, but I also don't call it
open source automatically because

173
00:11:46,050 --> 00:11:47,500
it depends on the license.

174
00:11:47,509 --> 00:11:52,360
Speaking generally, an open
source license is one that does

175
00:11:52,360 --> 00:11:54,260
not have use restrictions on it.

176
00:11:55,845 --> 00:12:00,125
And yet, as we'll talk later, there
are emerging some new kinds of

177
00:12:00,125 --> 00:12:04,765
licenses that, for AI safety reasons,
do actually attempt to restrict

178
00:12:04,984 --> 00:12:07,825
the allowable uses for a model.

179
00:12:09,164 --> 00:12:11,965
Okay, so let's get into the policy issues.

180
00:12:11,965 --> 00:12:15,865
It sounds companies can decide, open,
close, as they do with software.

181
00:12:16,365 --> 00:12:17,295
What's the concern?

182
00:12:17,765 --> 00:12:20,705
There are a number of concerns, but
I'd love to start with the benefits.

183
00:12:20,725 --> 00:12:23,925
Because really, the policy
question requires us.

184
00:12:24,270 --> 00:12:27,460
To weigh the benefits and the risks.

185
00:12:27,820 --> 00:12:31,660
And I think the benefits, although there
are some ways that open source software

186
00:12:31,660 --> 00:12:36,569
and open weights AI are not identical,
there are a lot of ways that they are

187
00:12:36,569 --> 00:12:38,810
similar, particularly in their benefits.

188
00:12:39,040 --> 00:12:41,789
And I break these benefits down
into three basic categories.

189
00:12:41,789 --> 00:12:47,330
One is simply distributing power,
whether in the market or in the

190
00:12:47,330 --> 00:12:50,130
culture, which I'll explain basically.

191
00:12:50,444 --> 00:12:56,694
As with open source software, we expect
and we currently see open weights AI

192
00:12:57,025 --> 00:13:03,845
being a strong competitive pressure
against open against closed providers, not

193
00:13:03,845 --> 00:13:08,634
least because it is free, we are seeing
very fast diffusion of the technology.

194
00:13:08,634 --> 00:13:12,505
We see very large enterprises
adopting it, including for internal

195
00:13:12,505 --> 00:13:14,455
uses like Dell and Wells Fargo.

196
00:13:14,455 --> 00:13:15,775
There's a great quote from Adele.

197
00:13:16,105 --> 00:13:17,215
Senior VP.

198
00:13:17,235 --> 00:13:21,835
That's basically like, why would we
pay for a general purpose model that

199
00:13:21,845 --> 00:13:23,855
does not know much about our company?

200
00:13:23,855 --> 00:13:26,464
And if we want to teach it that we
have to upload all of our private

201
00:13:26,465 --> 00:13:31,995
documents to their cloud and we have
to pay for it, or We could use this

202
00:13:32,005 --> 00:13:37,275
free software and create our own
bespoke model for our own purposes.

203
00:13:37,275 --> 00:13:41,615
That is more efficient, more tailored
to our needs, et cetera, et cetera.

204
00:13:41,854 --> 00:13:45,375
And so you're seeing that the
second big category of benefits

205
00:13:45,755 --> 00:13:48,974
is simply the catalyzing of.

206
00:13:49,414 --> 00:13:54,694
Innovation, not only innovation in AI, but
in all fields that can leverage the AI.

207
00:13:54,854 --> 00:13:59,865
The first LLMs that were
built were open research.

208
00:13:59,895 --> 00:14:04,555
They were openly published the Google,
the Google research scientists who

209
00:14:04,564 --> 00:14:07,115
created the first LLMs in 2017 ish.

210
00:14:07,444 --> 00:14:10,785
That is the root of the explosion
of AI innovation today is the fact

211
00:14:10,785 --> 00:14:12,495
that research happened in the open.

212
00:14:13,035 --> 00:14:14,325
Thanks to open development.

213
00:14:14,335 --> 00:14:17,295
You also see an explosion of developers.

214
00:14:17,605 --> 00:14:21,605
Taking open models and then
using them to build smaller, more

215
00:14:21,605 --> 00:14:23,495
efficient, customized models.

216
00:14:23,865 --> 00:14:27,949
Um, Build models that are small
enough to run locally rather than

217
00:14:27,949 --> 00:14:32,329
in the cloud, which has both privacy
benefits and environmental impact

218
00:14:32,329 --> 00:14:34,889
benefits and just bandwidth benefits.

219
00:14:34,909 --> 00:14:37,679
Like you don't have to spend
all that network infrastructure.

220
00:14:37,899 --> 00:14:42,560
Research around AI models also enables
a wide swath of security and safety

221
00:14:42,560 --> 00:14:44,700
research that can't happen effectively.

222
00:14:44,899 --> 00:14:49,420
with closed models and can enable
the faster development of including

223
00:14:49,430 --> 00:14:52,579
the faster development of tools
to detect and prevent bad things.

224
00:14:52,640 --> 00:14:55,770
There's another way in which it can
help security rather than hinder it.

225
00:14:56,130 --> 00:14:59,140
And then all the other
fields that leverage this AI,

226
00:14:59,159 --> 00:15:00,619
they will also move faster.

227
00:15:00,619 --> 00:15:04,679
And so there is a benefit of
simply speeding up the process of

228
00:15:04,679 --> 00:15:08,759
innovation and opening that process
to a lot more stakeholders as well.

229
00:15:09,279 --> 00:15:12,329
And then finally, there's the,
the transparency benefits.

230
00:15:12,735 --> 00:15:14,945
Which also brings security
and accountability benefits.

231
00:15:14,955 --> 00:15:19,275
Like with regular open source
software, with enough eyeballs, a lot

232
00:15:19,275 --> 00:15:24,005
of those problems will be discovered
and can be remedied because the

233
00:15:24,005 --> 00:15:27,415
whole world essentially is your red
team, the team that's testing it.

234
00:15:27,824 --> 00:15:30,614
There's a lot of research that's
happened that couldn't have happened

235
00:15:30,625 --> 00:15:34,625
without open source models around
guardrails, the guardrails that

236
00:15:34,625 --> 00:15:37,585
companies use to try to prevent bad
things from happening with their models.

237
00:15:37,915 --> 00:15:40,904
A lot of those, the challenges
with those guardrails have been

238
00:15:40,904 --> 00:15:42,345
discovered through open source.

239
00:15:42,790 --> 00:15:47,749
Models, um, develop in a way
that actually reflects on.

240
00:15:48,104 --> 00:15:51,275
The flaws of the closed models
as well, because they're similar

241
00:15:51,275 --> 00:15:55,314
architectures, the closed and the open
models, things that you learn about

242
00:15:55,314 --> 00:15:58,244
open models can be applied and you're
thinking about closed models as well.

243
00:15:58,275 --> 00:16:01,614
So it's a way of even getting more
transparency into the closed models,

244
00:16:01,755 --> 00:16:04,795
even though we don't actually have
literal transparency into the closed

245
00:16:04,795 --> 00:16:09,015
models and it's enabled research
around bias and discrimination that

246
00:16:09,015 --> 00:16:13,675
wouldn't have been possible without
access to the model and the weights.

247
00:16:13,974 --> 00:16:18,265
And one of the things we've discovered
in all that testing is, and this is

248
00:16:18,775 --> 00:16:22,335
Problematic and worrisome, but also
relevant to the analysis of whether open

249
00:16:22,335 --> 00:16:26,995
source is worth the risks, considering
the benefits is the guardrails that

250
00:16:26,995 --> 00:16:28,805
we are currently using to try to.

251
00:16:29,425 --> 00:16:31,954
Basically say no, have
a model say no to you.

252
00:16:31,954 --> 00:16:36,155
When you try to do something dangerous
or problematic, those are very fragile

253
00:16:36,454 --> 00:16:40,724
and they're fragile, whether or not
you are using open or closed models.

254
00:16:40,724 --> 00:16:45,324
And we wouldn't have known that without
research that was enabled by open models.

255
00:16:45,714 --> 00:16:48,574
This goes to an important concept
that we'll need to talk about when

256
00:16:48,574 --> 00:16:53,714
we talk about the risks, which is
what is the marginal risk of open

257
00:16:53,714 --> 00:16:56,015
source or open weights models.

258
00:16:56,655 --> 00:17:01,375
Compared to other technologies that are
available, whether it's the internet

259
00:17:01,375 --> 00:17:06,694
itself or closed models that are
also available to be used, because

260
00:17:06,704 --> 00:17:11,184
if there's not a big, or there, if
there's not a meaningful marginal

261
00:17:11,494 --> 00:17:15,594
risk that is differential risk between
those two things, there's no reason

262
00:17:15,625 --> 00:17:18,694
to target open source specifically.

263
00:17:18,990 --> 00:17:21,940
For restriction and we should
probably think about a different

264
00:17:21,940 --> 00:17:27,940
approach Okay Sounds wonderful
Why then are there concerns?

265
00:17:28,050 --> 00:17:32,440
Let me ask you in a more targeted
way Ntia national telecommunications

266
00:17:32,440 --> 00:17:35,530
and information administration in
the commerce department Launched

267
00:17:35,560 --> 00:17:41,610
this request for comment in early
2024 about open weights models.

268
00:17:41,610 --> 00:17:47,050
What, why did the government feel the
need to even ask whether these should

269
00:17:47,059 --> 00:17:49,839
be allowed to exist in a free form?

270
00:17:50,419 --> 00:17:53,039
Yeah, I think there are a number
of different categories of

271
00:17:53,040 --> 00:17:54,354
risk and a number of different.

272
00:17:54,735 --> 00:17:58,895
phases of concern that we've gone
through at this point in the past two

273
00:17:58,895 --> 00:18:05,484
years, since chat, GTP dropped for
real, I'd say the first category I

274
00:18:05,485 --> 00:18:08,944
would call emergent existential risks.

275
00:18:09,205 --> 00:18:14,965
I would say these were very
much at the forefront of.

276
00:18:15,770 --> 00:18:20,570
Initial thinking about the dangers of
AI because there is a community, we'll

277
00:18:20,570 --> 00:18:23,610
call them the AI safety community,
although there are plenty of other

278
00:18:23,639 --> 00:18:28,199
different types of AI safety folks
and stakeholders, there is a community

279
00:18:28,199 --> 00:18:30,839
around AI safety that essentially.

280
00:18:31,395 --> 00:18:36,875
Evolved over 20 years from work being
done in the Oxford philosophy department

281
00:18:36,895 --> 00:18:41,845
around potential existential risks
that humanity might face and one of

282
00:18:41,845 --> 00:18:46,035
those these people have theorized for
a while now is the possibility of some

283
00:18:46,354 --> 00:18:52,725
out of control super intelligent AI
that we lose control of and that does

284
00:18:52,725 --> 00:18:55,305
something catastrophic to harm humans.

285
00:18:56,080 --> 00:18:59,870
for its own reasons or because it
misunderstood its instructions.

286
00:19:00,280 --> 00:19:03,600
And I don't want to denigrate the
people who are raising this concern.

287
00:19:03,959 --> 00:19:07,979
Speaking generally, I think these
are very intelligent, very well

288
00:19:07,979 --> 00:19:13,279
intentioned, sincere people who want
to see these concerns addressed and

289
00:19:13,279 --> 00:19:15,210
don't want to see AI destroy the world.

290
00:19:17,395 --> 00:19:23,015
The more likely threat than the sort
of vague super AI is the possibility

291
00:19:23,045 --> 00:19:28,275
that these systems, as they get more
intelligent, could significantly

292
00:19:28,315 --> 00:19:32,874
aid adversaries, whether state
adversaries or non state adversaries,

293
00:19:33,174 --> 00:19:37,319
in developing chemical, biological,
radiological, or nuclear weapons.

294
00:19:38,040 --> 00:19:41,760
or cyber weapons, hacking,
hacking tools and whatnot, and

295
00:19:41,790 --> 00:19:43,510
automated attack approaches.

296
00:19:44,010 --> 00:19:48,650
We will talk at more length about
that because it ties in to the NTIA

297
00:19:48,840 --> 00:19:53,429
process, but really the key issue
there, as is the key issue generally,

298
00:19:53,479 --> 00:19:55,439
is there actually a marginal risk?

299
00:19:55,870 --> 00:19:59,769
Does open source pose a greater risk of
those things than other technologies?

300
00:19:59,809 --> 00:20:01,719
And I'm not going to spoil that.

301
00:20:02,530 --> 00:20:03,889
I'll tell you how that turned out.

302
00:20:04,260 --> 00:20:09,210
A third category is simply content
issues of harmful or problematic

303
00:20:09,210 --> 00:20:15,109
or illegal content, particularly of
concern are deep faked child sexual

304
00:20:15,109 --> 00:20:21,890
abuse material, virtual CSAM or deep
faked NCII non consensual intimate

305
00:20:21,909 --> 00:20:23,989
imagery, often called revenge porn.

306
00:20:24,489 --> 00:20:27,589
And this is going to be a
problem and is a problem.

307
00:20:27,920 --> 00:20:30,930
But again, the marginal risk issue
is a question that needs to be

308
00:20:30,930 --> 00:20:32,230
asked, and we'll ask it shortly.

309
00:20:33,310 --> 00:20:36,500
Finally, the fourth category, and
where I think most of the concern

310
00:20:36,500 --> 00:20:40,290
is settling now, and which I'm sure
we'll talk about later, is China.

311
00:20:41,230 --> 00:20:42,500
China specifically.

312
00:20:42,709 --> 00:20:48,259
What do we do about Chinese
competition in the AI field?

313
00:20:48,880 --> 00:20:52,900
How do we deal with China as
a national security competitor

314
00:20:53,220 --> 00:20:55,180
in regard to AI as well?

315
00:20:55,430 --> 00:20:57,450
And I think we are at a point now.

316
00:20:57,754 --> 00:21:03,395
Where that's probably the most relevant
consideration for policymakers in the U.

317
00:21:03,395 --> 00:21:03,784
S.

318
00:21:04,195 --> 00:21:07,764
As they look toward possible
restrictions on open source now.

319
00:21:08,314 --> 00:21:11,554
But, so that, that's the sort
of range of risks I think

320
00:21:11,574 --> 00:21:13,405
people are most concerned about.

321
00:21:14,205 --> 00:21:17,224
Okay, let's drill down on some of those.

322
00:21:17,224 --> 00:21:19,735
I, I agree we can put
aside the existential risk.

323
00:21:19,800 --> 00:21:24,770
But with regard to marginal risk,
you made the point that, that guard

324
00:21:24,770 --> 00:21:30,180
rails on these foundation models are
fragile, but isn't it true that it's

325
00:21:30,400 --> 00:21:35,050
substantially easier to remove a guard
rail if you have access to the weight.

326
00:21:35,050 --> 00:21:40,060
So if it's a, an open model, then it's
fairly trivial for someone to take a,

327
00:21:40,170 --> 00:21:44,410
that's a hostile actor to take out those
guard rails, much more challenging.

328
00:21:44,420 --> 00:21:46,339
If the company keeps
control over the model.

329
00:21:46,850 --> 00:21:51,020
It depends the you may
have seen chat GTPs.

330
00:21:53,304 --> 00:21:54,675
model released last week.

331
00:21:54,695 --> 00:21:59,475
And then within minutes on Twitter,
people sharing their, their attacks,

332
00:21:59,605 --> 00:22:04,204
their prompts that were able to get
around certain content guardrails.

333
00:22:04,435 --> 00:22:08,645
The Taylor Swift sexually explicit
deep fakes that we saw making the

334
00:22:08,645 --> 00:22:13,534
rounds earlier in the summer, those
were generated by Microsoft's.

335
00:22:13,715 --> 00:22:19,834
Closed image creator, but yes, open
source code, open source models are

336
00:22:20,314 --> 00:22:25,715
more easily modifiable in a wide range
of ways, including by bad actors.

337
00:22:26,055 --> 00:22:30,215
But then that, that, that leads to the
question of what about other technologies?

338
00:22:30,215 --> 00:22:31,344
How does it compare to those?

339
00:22:31,344 --> 00:22:34,465
And that's something that the NTIA
got into, but to set up the NTIA

340
00:22:34,465 --> 00:22:41,475
process a This came up in the context
of the Biden AI executive order that

341
00:22:41,475 --> 00:22:48,515
came out October 30th, I want to say
last year, and which called on NTIA

342
00:22:48,815 --> 00:22:52,625
to do a report on open weights AI.

343
00:22:53,074 --> 00:22:59,665
This was prompted as best we can tell by
there being staff at the National Security

344
00:22:59,665 --> 00:23:02,815
Council of the White House and staff
at the Bureau of Industry and Security,

345
00:23:02,815 --> 00:23:07,754
the export control guys at the Commerce
Department who were concerned particularly

346
00:23:07,754 --> 00:23:09,635
about the CBRN and cyber risks.

347
00:23:09,800 --> 00:23:16,060
And we're considering the possibility
or looking at the possibility

348
00:23:16,100 --> 00:23:19,340
of putting restrictions on the
publication of open weights because

349
00:23:19,340 --> 00:23:20,870
of those national security concerns.

350
00:23:21,170 --> 00:23:24,739
It seems that somewhat cooler
heads prevailed in pushing

351
00:23:25,149 --> 00:23:27,269
instead pushing for that.

352
00:23:27,760 --> 00:23:31,899
Not to be decided in the EO, but rather
to delegate to NTIA, the National

353
00:23:31,899 --> 00:23:35,909
Telecommunications Information and
Information Administration, essentially

354
00:23:35,909 --> 00:23:38,580
the president's lawyer, the president,
not the president's lawyers, the

355
00:23:38,580 --> 00:23:44,660
president's technology advisors, um, to
do a more sustained study of the issue.

356
00:23:44,970 --> 00:23:49,679
The issue in particular being what are
the benefits of open weights models?

357
00:23:50,030 --> 00:23:53,830
What are the risks of open
weights models based on that?

358
00:23:53,840 --> 00:23:57,920
What are potential policy approaches
the president should take?

359
00:23:58,360 --> 00:24:03,280
Or not and and so that was the process
and they put out the call for comment

360
00:24:03,300 --> 00:24:07,370
earlier this year A bunch of people
filed comments including us and i'll

361
00:24:07,370 --> 00:24:11,090
talk about that And then they finally
released a report I guess a couple of

362
00:24:11,090 --> 00:24:15,189
months ago And so let's talk about that
and i'll talk about the comments that

363
00:24:15,189 --> 00:24:17,860
cdt filed first a not so humble brag.

364
00:24:17,860 --> 00:24:21,825
I'm proud to say This is a dubious
achievement, but our comments were the

365
00:24:21,865 --> 00:24:25,875
longest and most detailed comments filed
in the proceeding, which had hundreds upon

366
00:24:25,885 --> 00:24:31,684
hundreds of comments, more salient, our
comments were cited in the report from

367
00:24:31,684 --> 00:24:34,625
the NTIA more than any other comments.

368
00:24:34,835 --> 00:24:39,335
And the only other single source that
was cited more than our comments was a

369
00:24:39,375 --> 00:24:43,015
paper that we coauthored with academics
at Stanford and Princeton focused on

370
00:24:43,015 --> 00:24:45,945
defining marginal risk around open models.

371
00:24:46,165 --> 00:24:50,925
I'm glad to say, I feel that we had a
very positive influence on this process.

372
00:24:51,165 --> 00:24:56,075
We also worked closely with NTIA to ensure
they heard from a variety of other civil

373
00:24:56,075 --> 00:25:00,205
society voices from civil rights groups
to civil liberties groups to open source

374
00:25:00,205 --> 00:25:06,215
advocates and everything in between, but
the gist of what we were saying in the

375
00:25:06,215 --> 00:25:09,444
paper I mentioned in our comments is.

376
00:25:10,840 --> 00:25:14,670
We are not ruling out categorically
the possibility that some restrictions

377
00:25:14,670 --> 00:25:21,319
on open weights may ultimately be
necessary, but at this point at, with the

378
00:25:21,320 --> 00:25:27,280
capabilities we see online now and that
are coming shortly, there is not enough

379
00:25:27,320 --> 00:25:30,700
evidence to support those restrictions.

380
00:25:31,290 --> 00:25:34,480
And by, by that, not enough evidence
that there's actually a significant

381
00:25:34,520 --> 00:25:37,050
marginal risk from open models.

382
00:25:37,440 --> 00:25:39,770
This is probably clearest.

383
00:25:40,190 --> 00:25:47,410
In the context of the like biological and
nuclear and cyber conversation there by

384
00:25:47,410 --> 00:25:53,990
the time of our comments, there had been
several academic papers and papers from

385
00:25:54,319 --> 00:26:00,170
security think tanks like Rand research by
Microsoft and open AI research where they

386
00:26:00,700 --> 00:26:05,330
basically set up teams to compete on who
could come up with the best evil terrorist

387
00:26:05,340 --> 00:26:12,220
plan around biological weapons and
gave One team access to the best models

388
00:26:12,720 --> 00:26:14,840
and one team access to the internet.

389
00:26:15,580 --> 00:26:20,409
And what they found was there was very
little difference in the capacity of

390
00:26:20,409 --> 00:26:26,470
those teams to plan an attack, which,
which stands to reason because what the

391
00:26:26,480 --> 00:26:29,279
models know is what they've seen online.

392
00:26:29,720 --> 00:26:34,169
And so it may give some
incremental additional aid in.

393
00:26:34,835 --> 00:26:37,985
in collecting some of your sources
when you're looking for information

394
00:26:37,985 --> 00:26:41,655
to help plot a plot, but it really
didn't make a huge difference.

395
00:26:42,005 --> 00:26:49,000
Similarly, There was research on how much
models, which certain models can also

396
00:26:49,000 --> 00:26:53,379
help instead of just speaking in language,
they can also code, they can create

397
00:26:53,379 --> 00:26:58,979
software, and so there were tests about
how effective they were in helping create

398
00:26:59,590 --> 00:27:02,409
penetration tools and other attack tools.

399
00:27:02,839 --> 00:27:08,519
And again, they found only a incremental
difference, not really a major difference.

400
00:27:08,970 --> 00:27:12,910
between what a coder could do without a
model and what a coder could do with a

401
00:27:12,910 --> 00:27:15,850
model in terms of offensive cyber stuff.

402
00:27:16,230 --> 00:27:23,189
It's also worth noting, as with the case
with open source software, that capability

403
00:27:23,189 --> 00:27:28,880
to, to create tools also is now in the
hands of defenders as well as attackers.

404
00:27:28,890 --> 00:27:33,070
And as we've seen in the open source
context, that has tended to ultimately

405
00:27:33,070 --> 00:27:35,110
benefit defenders more than attackers.

406
00:27:35,530 --> 00:27:40,020
And so NTIA, which was primarily
focused on those as the most realistic,

407
00:27:40,515 --> 00:27:43,765
National security related threats
that they were asked to assess

408
00:27:44,185 --> 00:27:48,485
essentially concluded We don't see
a significant marginal risk here.

409
00:27:49,324 --> 00:27:53,264
That doesn't mean there might
not be one at some point.

410
00:27:53,785 --> 00:27:58,584
Therefore, we do endorse
continued monitoring.

411
00:27:58,794 --> 00:28:01,984
I guess what a doctor would call watchful
waiting, like it's not something that

412
00:28:01,985 --> 00:28:05,364
requires intervention now, but something
that we need to be watching closely,

413
00:28:05,664 --> 00:28:11,034
including recommending continued
investment in the development of best

414
00:28:11,034 --> 00:28:13,244
practices and standards around safety.

415
00:28:13,770 --> 00:28:18,100
As well as market surveillance
and other measures to see how

416
00:28:18,100 --> 00:28:19,639
these models are being used.

417
00:28:20,050 --> 00:28:23,740
Because so much of the risk is actually
going to turn on what context it is

418
00:28:23,750 --> 00:28:25,550
being used and by what types of users.

419
00:28:26,009 --> 00:28:30,959
And basically advise the president
to hold off on restrictive measures.

420
00:28:31,380 --> 00:28:35,319
And I think this was the right
decision in a number of ways.

421
00:28:36,910 --> 00:28:42,389
But especially by analogy to a similar
fight over tech policy that we had

422
00:28:42,389 --> 00:28:45,860
in the 90s, that we haven't talked
about yet, but that has to be talked

423
00:28:45,870 --> 00:28:49,889
about in this context, which is the
fight over open source encryption

424
00:28:50,510 --> 00:28:53,000
code in the 90s and early 2000s.

425
00:28:53,320 --> 00:28:57,190
I think we have a very
similar issue with AI today.

426
00:28:57,200 --> 00:29:02,140
And this is now I'll talk about China,
because it's a very similar situation

427
00:29:02,140 --> 00:29:03,850
where in the encryption case, it was.

428
00:29:04,254 --> 00:29:08,575
We are afraid our adversaries
will misuse this technology and

429
00:29:08,575 --> 00:29:11,495
therefore we want to try to limit
the spread of this technology.

430
00:29:11,495 --> 00:29:14,605
But the technology we're
talking about is literally bits.

431
00:29:14,815 --> 00:29:18,095
It's not a physical thing that
we can physically constrain.

432
00:29:18,425 --> 00:29:23,374
And so it's really hard to prevent the
spread of software on a global internet.

433
00:29:23,804 --> 00:29:27,504
And that's as true today, if
not more than it was then.

434
00:29:27,884 --> 00:29:33,465
And there are those that are concerned
that if we are doing open source AI models

435
00:29:33,504 --> 00:29:39,705
that We are giving away our IP to our
competitors in China, that we are giving

436
00:29:39,705 --> 00:29:44,744
away software that could be integrated
into military operations by China.

437
00:29:46,044 --> 00:29:50,674
Both of those are true in the sense that
certainly there will be stakeholders in

438
00:29:50,674 --> 00:29:53,015
China that make use of this software, but.

439
00:29:54,405 --> 00:29:58,345
Even if our software is not available,
they will make use of similar software,

440
00:29:58,345 --> 00:30:04,465
whether it comes from France or the United
Arab Emirates or China itself, where

441
00:30:04,465 --> 00:30:10,514
there are plenty of large models being
created today without reliance on U.

442
00:30:10,514 --> 00:30:10,654
S.

443
00:30:10,664 --> 00:30:11,314
technology.

444
00:30:11,745 --> 00:30:15,945
So I think it's a question of, yes,
we could attempt to restrict this.

445
00:30:16,685 --> 00:30:18,665
We would probably fail.

446
00:30:18,675 --> 00:30:22,634
It would probably leak anyway,
but even if we did not fail, it

447
00:30:22,635 --> 00:30:24,335
likely wouldn't be effective.

448
00:30:24,835 --> 00:30:29,475
in preventing China from getting
comparable technology elsewhere.

449
00:30:29,935 --> 00:30:36,985
Meanwhile, we will have slowed down our
own innovation cycle, slowed down the

450
00:30:36,985 --> 00:30:41,105
diffusion of the technology generally,
and the many economic benefits and

451
00:30:41,105 --> 00:30:46,465
other benefits that will accrue from
that, and have hindered our ability

452
00:30:46,465 --> 00:30:49,135
to compete in the global marketplace.

453
00:30:49,585 --> 00:30:54,595
Around this kind of software, which I
think is especially important to consider

454
00:30:54,605 --> 00:30:58,305
when you think about China's geopolitical
position right now and things like

455
00:30:58,305 --> 00:31:03,455
the Belt and Road Initiative, which is
its massive global attempt to provide

456
00:31:03,475 --> 00:31:08,685
infrastructure, including communications
infrastructure to the global South,

457
00:31:08,695 --> 00:31:13,625
to Africa, to South America, to the
Middle East as a way of consolidating

458
00:31:13,625 --> 00:31:19,425
geopolitical power, our having a
robust open ecosystem of AI models.

459
00:31:19,875 --> 00:31:26,085
Is probably our best shot at
preventing China from dominating the

460
00:31:26,095 --> 00:31:29,775
AI market in those global markets.

461
00:31:30,254 --> 00:31:33,964
And they're actually at a unique
disadvantage compared to our open

462
00:31:33,964 --> 00:31:37,085
models because, because they are China.

463
00:31:37,355 --> 00:31:42,295
The CCP has already passed
regulations that basically enforce

464
00:31:42,425 --> 00:31:45,985
ideological purity on Chinese models.

465
00:31:46,215 --> 00:31:52,185
Basically, censorship in line with
The party line in china, which makes

466
00:31:52,185 --> 00:31:56,635
them necessarily much less useful to
a lot of stakeholders So if we can

467
00:31:56,635 --> 00:32:04,955
compete on a global stage With free,
powerful, customizable AI software,

468
00:32:05,435 --> 00:32:10,615
and China is competing with less good
software, then that actually strengthens

469
00:32:10,615 --> 00:32:14,984
our geopolitical position in regard
to China, rather than weakening it.

470
00:32:15,775 --> 00:32:21,290
Would be Our argument and the argument
at this point of a number of relatively

471
00:32:21,290 --> 00:32:26,729
conservative stakeholders as well as
your center left types like me Okay,

472
00:32:26,750 --> 00:32:30,649
there's a lot more that I would love
to get into this with you about but

473
00:32:30,670 --> 00:32:34,250
we have limited time I don't want to
just ask one or two more questions.

474
00:32:34,359 --> 00:32:40,270
Sure One is that the company used
to work for, Meta, is the only one

475
00:32:40,270 --> 00:32:44,719
of the major frontier AI labs that,
that seems to be pushing forward with

476
00:32:44,720 --> 00:32:49,709
releasing its most powerful models,
the LLAMA models as open weights.

477
00:32:50,100 --> 00:32:53,890
So can you understand, of course,
you, you are not speaking for them.

478
00:32:54,119 --> 00:32:55,719
Can you talk a little bit about.

479
00:32:56,165 --> 00:33:01,025
Maybe why it seems like they've taken
that direction and how is that going

480
00:33:01,025 --> 00:33:05,714
to affect this debate that it's not now
just an issue about AI safety, but it

481
00:33:05,714 --> 00:33:10,115
plays into all of the competitive and
other debates in the technology industry.

482
00:33:10,435 --> 00:33:10,865
Sure.

483
00:33:10,925 --> 00:33:13,834
Yeah, I'm not going to attempt
to divine the mind of Mark

484
00:33:13,834 --> 00:33:15,805
Zuckerberg or say anything.

485
00:33:16,285 --> 00:33:19,785
Confidential that I may or may not
have learned inside of Meta, but I

486
00:33:19,795 --> 00:33:23,925
can note some things that other people
have said and some obvious facts.

487
00:33:24,425 --> 00:33:30,444
I think one of them is simple is
the simple fact that Facebook now

488
00:33:30,444 --> 00:33:34,755
Meta doesn't have a cloud services
division through which it wants

489
00:33:34,755 --> 00:33:37,074
to sell access to closed models.

490
00:33:37,564 --> 00:33:40,204
It's it's whole business
model is different.

491
00:33:40,224 --> 00:33:42,414
It offers consumer products.

492
00:33:42,935 --> 00:33:48,775
To consumers that rely on AI,
but the them dominating the AI

493
00:33:49,185 --> 00:33:51,765
model market is not their goal.

494
00:33:52,124 --> 00:33:55,914
As one could argue is the goal of
your Google's, your Microsoft's,

495
00:33:55,914 --> 00:33:57,675
your Amazon's, your open AI's.

496
00:33:58,195 --> 00:34:00,174
So they just have different incentives.

497
00:34:00,680 --> 00:34:07,230
One of those incentives is by opening
their largest model, which is also,

498
00:34:07,400 --> 00:34:12,810
they've said the model that is going to
be the primary basis of a lot of features

499
00:34:12,810 --> 00:34:17,630
in their products, they can get the
traditional benefits of open development.

500
00:34:17,660 --> 00:34:21,670
They can get those millions of
eyes on their, on their model.

501
00:34:21,680 --> 00:34:26,300
They can benefit from seeing
what customizations people

502
00:34:26,300 --> 00:34:27,770
are doing on platforms like.

503
00:34:28,020 --> 00:34:32,960
GitHub or HuggingFace and integrating the
good ones into their production model.

504
00:34:33,300 --> 00:34:38,520
And then there's also simply finally
the obvious fact that it is what

505
00:34:38,520 --> 00:34:41,100
they meta when ChachiTP dropped.

506
00:34:41,440 --> 00:34:46,719
Meta had long been doing open
source releases of research models.

507
00:34:46,950 --> 00:34:51,630
Some may recall the dubious release
of Galactica shortly before ChachiTP,

508
00:34:51,640 --> 00:34:56,160
which was a, an LLM trained on and
focused on scientific literature.

509
00:34:56,535 --> 00:34:59,275
That spouted nonsense and was
taken down pretty quickly.

510
00:34:59,595 --> 00:35:05,654
And so you already had this large research
investment in developing LLM technology

511
00:35:05,654 --> 00:35:07,985
that wasn't necessarily being productized.

512
00:35:08,555 --> 00:35:10,175
And there's the obvious fact that if.

513
00:35:10,660 --> 00:35:14,950
Meta threw that out there while
open AI and Microsoft were trying

514
00:35:14,950 --> 00:35:16,890
to consolidate a closed AI position.

515
00:35:17,170 --> 00:35:19,170
It would slow down those competitors.

516
00:35:19,470 --> 00:35:22,760
Many of your listeners may have
heard of the memo inside of Google.

517
00:35:23,090 --> 00:35:27,809
I believe it was called open AI has no
moat and neither do we in the sense of

518
00:35:27,809 --> 00:35:30,650
a competitive moat against open source.

519
00:35:31,070 --> 00:35:31,810
And I think.

520
00:35:32,325 --> 00:35:35,575
People inside of Facebook
probably followed a similar logic.

521
00:35:35,635 --> 00:35:37,045
I am neither confirming or denying.

522
00:35:37,045 --> 00:35:40,344
And in fact, I was on parental leave
for much of this decision making, but

523
00:35:40,345 --> 00:35:46,505
recognized that offering this would not
only create a lot of benefit and use

524
00:35:46,505 --> 00:35:51,294
for a lot of people, but also slow the
adoption of closed models offered by

525
00:35:51,295 --> 00:35:57,215
their competitors, who again, Meta doesn't
offer a cloud service, but it does offer

526
00:35:57,215 --> 00:36:00,275
a lot of other services that are in
competition with a lot of these companies.

527
00:36:00,635 --> 00:36:04,954
And so I think as a number of
commentators have noted, that was a

528
00:36:04,954 --> 00:36:09,494
clear benefit to Facebook if they could
throw an obstacle in front of their

529
00:36:09,494 --> 00:36:11,694
fast moving opponents in the market.

530
00:36:12,615 --> 00:36:16,065
As I make the distinction between
open weights and open source, Meta is

531
00:36:16,065 --> 00:36:19,285
a good example of why we need to be
really careful about our terminology

532
00:36:19,285 --> 00:36:24,795
there, because Mark Zuckerberg loves
to call LLAMA open source and use

533
00:36:24,795 --> 00:36:29,214
that terminology, but the license
that Meta is licensing under is

534
00:36:29,224 --> 00:36:31,579
definitely not an open source license.

535
00:36:32,200 --> 00:36:36,180
In some ways that are arguably good
and in some ways that are arguably bad.

536
00:36:36,390 --> 00:36:41,440
Most notably, their original
license forbade you from building

537
00:36:41,440 --> 00:36:44,500
a model using outputs from their
model, which is actually a very

538
00:36:44,500 --> 00:36:48,279
powerful way of building smaller,
more condensed versions of models.

539
00:36:48,720 --> 00:36:49,989
Thankfully, they fixed that.

540
00:36:49,999 --> 00:36:58,590
But now if you do that, your model has
to be Marked as built with llama and you

541
00:36:58,590 --> 00:37:03,530
don't typically see that kind of branding
requirement in the open source community

542
00:37:03,750 --> 00:37:07,710
And so I'd love for them to get rid
of that and then also as has been much

543
00:37:07,710 --> 00:37:13,479
commented on they also have a clause that
prohibits any company with more than 700

544
00:37:13,519 --> 00:37:19,975
million average monthly users To use the
model this is to basically prevent google

545
00:37:19,975 --> 00:37:23,845
and microsoft from using the model for
free And if they want to use it, they'll

546
00:37:23,845 --> 00:37:28,404
have to go to to meta to license it again
Like that's not the sort of thing you'd

547
00:37:28,404 --> 00:37:33,664
see in a regular open source license
carving out particular users At least

548
00:37:33,664 --> 00:37:38,790
it is limited to Companies that had that
level of users at the time of the license.

549
00:37:38,800 --> 00:37:43,930
So it's not like you could build your
company on LLAMA, ultimately get to 750

550
00:37:43,950 --> 00:37:48,269
and then suddenly have the bottom pulled
out of your business, but it's still fair

551
00:37:48,269 --> 00:37:53,229
to say not anti competitive and something
that, that especially if they want to

552
00:37:53,229 --> 00:37:57,300
keep calling it open source, they need
to take out of their license or they

553
00:37:57,300 --> 00:37:58,480
need to stop calling it open source.

554
00:38:00,190 --> 00:38:03,450
All right, there's a lot more that we
could talk about here Obviously the

555
00:38:03,450 --> 00:38:07,410
ntia report will not be the last word,
but we're going to need to wrap up for

556
00:38:07,410 --> 00:38:11,239
time kevin Thank you so much for going
through and in so much detail this really

557
00:38:11,239 --> 00:38:13,520
important set of issues with us Anytime.

558
00:38:13,870 --> 00:38:14,480
It was a pleasure.

559
00:38:15,030 --> 00:38:15,660
Great to have you.

560
00:38:17,590 --> 00:38:19,500
This has been The Road to Accountable AI.

561
00:38:19,700 --> 00:38:22,360
If you like what you're hearing,
please give us a good review and

562
00:38:22,360 --> 00:38:25,770
check out my sub stack for more
insights on AI accountability.

563
00:38:26,140 --> 00:38:26,889
Thank you for listening.

564
00:38:35,069 --> 00:38:35,950
This is Kevin Werbeck.

565
00:38:36,309 --> 00:38:39,805
If you want to go deeper, check On AI
governance, trust, and responsibility

566
00:38:39,805 --> 00:38:43,245
with me and other distinguished
faculty of the world's top business

567
00:38:43,255 --> 00:38:47,214
school, sign up for the next cohort of
Wharton's strategies for accountable

568
00:38:47,245 --> 00:38:51,715
AI online executive education
program, featuring live interaction

569
00:38:51,715 --> 00:38:55,615
with faculty, expert interviews, and
custom designed asynchronous content.

570
00:38:56,194 --> 00:38:59,295
Join fellow business leaders to
learn valuable skills you can

571
00:38:59,305 --> 00:39:01,015
put to work in your organization.

572
00:39:01,765 --> 00:39:03,135
Visit execed.

573
00:39:03,145 --> 00:39:03,595
wharton.

574
00:39:03,595 --> 00:39:04,565
upenn.

575
00:39:04,565 --> 00:39:08,455
edu slash ACAI for full details.

576
00:39:08,995 --> 00:39:09,805
I hope to see you there.

