1
00:00:01,967 --> 00:00:03,257
Elham Tabasi, welcome.

2
00:00:03,257 --> 00:00:05,045
Glad to have you on the podcast.

3
00:00:05,262 --> 00:00:06,362
Thanks for having me.

4
00:00:06,362 --> 00:00:07,686
Glad to be here.

5
00:00:08,751 --> 00:00:15,551
But first, why don't you tell us what NIST
is, because it's an important agency, but

6
00:00:15,551 --> 00:00:20,771
one that maybe people are not necessarily
familiar with, even if they know something

7
00:00:20,771 --> 00:00:23,839
about the US government and AI policy.

8
00:00:24,782 --> 00:00:25,792
Oh, yeah, happy to.

9
00:00:25,792 --> 00:00:28,002
And that's a great starting point.

10
00:00:28,162 --> 00:00:33,142
So NIST, or National Institute of
Standards and Technology, is a non

11
00:00:33,142 --> 00:00:34,402
-regulatory agency.

12
00:00:34,402 --> 00:00:36,962
We are under the Department of Commerce.

13
00:00:37,002 --> 00:00:43,122
And NIST was established back in 1902 with
a unique mission to advance US innovation

14
00:00:43,122 --> 00:00:45,802
and industrial competitiveness.

15
00:00:46,482 --> 00:00:50,382
Here at NIST, we have a very broad
portfolio of research.

16
00:00:50,382 --> 00:00:55,102
from building the most accurate atomic
clock to cybersecurity to artificial

17
00:00:55,102 --> 00:01:00,002
intelligence, advanced communications, but
most importantly, a long tradition of

18
00:01:00,002 --> 00:01:02,382
cultivating trust in technology.

19
00:01:02,862 --> 00:01:06,621
And we do that by advancing measurement
science and standards, measurement science

20
00:01:06,621 --> 00:01:12,682
and standards that makes technology more
reliable, secure, private, fair, or in

21
00:01:12,682 --> 00:01:14,132
other words, more trustworthy.

22
00:01:14,132 --> 00:01:17,310
And that's exactly what we have been doing
in the space of AI.

23
00:01:18,273 --> 00:01:19,073
Absolutely.

24
00:01:19,073 --> 00:01:23,273
And so you were one of the architects of
the NIST AI Risk Management framework,

25
00:01:23,273 --> 00:01:26,973
which came out in January of 2023.

26
00:01:27,593 --> 00:01:31,193
Other than the general goal of trust and
technology that you just described, where

27
00:01:31,193 --> 00:01:32,603
did that project come from?

28
00:01:32,603 --> 00:01:34,055
How did that come about?

29
00:01:34,894 --> 00:01:35,934
Yeah.

30
00:01:36,254 --> 00:01:45,534
So AIRMF that was released a little over a
year ago, it was directed by congressional

31
00:01:45,534 --> 00:01:46,674
mandate.

32
00:01:46,674 --> 00:01:52,974
It is a voluntary framework for managing
the risk of AI in a flexible, structured,

33
00:01:52,974 --> 00:01:55,554
and measurable way.

34
00:01:56,774 --> 00:01:58,990
The whole point is that

35
00:01:58,990 --> 00:02:04,310
We harness the power of AI for good and AI
can serve all people in a beneficial, fair

36
00:02:04,310 --> 00:02:06,730
and equitable manner.

37
00:02:07,750 --> 00:02:11,030
These three attributes that I talked
about, flexible, structured and

38
00:02:11,030 --> 00:02:18,370
measurable, we wanted to be flexible to be
able to keep up the pace with this rapid

39
00:02:18,370 --> 00:02:23,662
change of the technology, but also so that
the different organizations...

40
00:02:23,662 --> 00:02:27,262
with different resources be able to
implement that.

41
00:02:27,262 --> 00:02:33,162
We didn't want to be a heavy thing that
only the big organizations that have AI

42
00:02:33,162 --> 00:02:38,002
governance or a lot of people doing that
be able to operationalize it.

43
00:02:38,362 --> 00:02:44,202
Structured in the terms that when we start
working on AI RMF, the concept of the

44
00:02:44,202 --> 00:02:49,082
trustworthiness, the question of what
constitutes trust were not quite answered

45
00:02:49,082 --> 00:02:52,558
when there has been a lot of really high
value

46
00:02:52,558 --> 00:02:59,188
high -level value -based documents that
talk about these things.

47
00:02:59,188 --> 00:03:06,638
None of them try to get those principles
into sort of the building blocks that the

48
00:03:06,638 --> 00:03:09,398
developers and the employer of the
technology can use that.

49
00:03:09,398 --> 00:03:13,818
So working with the community, we try to
come up with unpacking the concept of the

50
00:03:13,818 --> 00:03:18,378
trustworthiness, and come up with the
components of the AI system that's

51
00:03:18,378 --> 00:03:22,350
trustworthy, which those characteristics
are.

52
00:03:22,350 --> 00:03:28,570
valid and reliable, safe, secure and
resilient, privacy -enhanced, explainable

53
00:03:28,570 --> 00:03:34,790
and interpretable, fair with harmful bias,
managed and transparent and accountable.

54
00:03:34,870 --> 00:03:40,470
Bring the community on a shared
understanding around each of them to get

55
00:03:40,470 --> 00:03:46,450
to the next attribute of the measurable to
have a good understanding of what we want

56
00:03:46,450 --> 00:03:48,710
to measure to be able to measure them.

57
00:03:48,710 --> 00:03:52,014
So the importance of being able to...

58
00:03:52,014 --> 00:03:58,554
define trustworthiness into technical and
socio -technical terms that developers of

59
00:03:58,554 --> 00:04:02,174
the technology, deployers of the
technologies, users of the technology can

60
00:04:02,174 --> 00:04:06,134
understand, and then build a community
together to come up with the scientific

61
00:04:06,134 --> 00:04:11,094
underpinning for how to measure those
because if you're really serious in the

62
00:04:11,094 --> 00:04:14,534
business of improving things, first we
have to be able to measure them.

63
00:04:14,534 --> 00:04:17,492
We can't improve things that we cannot
measure.

64
00:04:18,543 --> 00:04:21,203
So I want to get back to that question
about measurement because it is really

65
00:04:21,203 --> 00:04:21,723
critical.

66
00:04:21,723 --> 00:04:25,703
But first, let me just ask you, that's a
big challenge talking about

67
00:04:25,703 --> 00:04:28,083
trustworthiness around AI and all those
principles you described.

68
00:04:28,083 --> 00:04:29,363
They make sense.

69
00:04:29,363 --> 00:04:32,763
But how did you and the team get your arms
around that?

70
00:04:32,763 --> 00:04:37,103
How to decide what kinds of aspects to
focus on and how to operationalize these

71
00:04:37,103 --> 00:04:38,287
broad principles?

72
00:04:39,086 --> 00:04:40,326
Great question.

73
00:04:40,326 --> 00:04:46,526
So how we do our work in this case and
also in many other cases is to work

74
00:04:46,526 --> 00:04:49,586
collaboratively with public and private
stakeholders.

75
00:04:49,586 --> 00:04:54,666
So basically everything that we do, we
start with a formal process of request for

76
00:04:54,666 --> 00:04:55,245
information.

77
00:04:55,245 --> 00:04:58,486
That's where the government put an
announcement into a federal registered

78
00:04:58,486 --> 00:05:03,326
notice, announce to the whole public that
we're doing this and ask for comments.

79
00:05:03,326 --> 00:05:08,846
Then we work with the community to do
rounds of workshops, listening sessions.

80
00:05:08,846 --> 00:05:13,346
to get answers to all of these questions.

81
00:05:13,906 --> 00:05:22,166
And in doing so, we very early on realized
that we should definitely listen to a

82
00:05:22,166 --> 00:05:25,586
community with the expertise that are
building the technology or designing the

83
00:05:25,586 --> 00:05:29,986
technology, those that have the expertise
in computer science, mathematics,

84
00:05:29,986 --> 00:05:34,416
statistician, but also the community that
study the impact of the technology,

85
00:05:34,416 --> 00:05:38,286
psychologists, sociologists, philosophers.

86
00:05:38,286 --> 00:05:43,826
but also the community that are going to
be in charge of either top down or bottom

87
00:05:43,826 --> 00:05:48,326
up, try to operationalize and implement it
and figure out how it basically fits into

88
00:05:48,326 --> 00:05:54,006
the bigger risk management of the
organization, so business groups.

89
00:05:54,265 --> 00:06:00,066
And we also did listening sessions with
the AI users, one of our best listening

90
00:06:00,066 --> 00:06:05,346
sessions when we went to community
colleges, when we went to minority serving

91
00:06:05,346 --> 00:06:07,502
institutions and talked with the students.

92
00:06:07,502 --> 00:06:14,642
So by getting a lot of input from the
community, input from a very diverse sets

93
00:06:14,642 --> 00:06:23,561
of backgrounds and expertise, we came to
where we are in terms of the

94
00:06:23,561 --> 00:06:26,528
trustworthiness characteristics, but also
the whole document.

95
00:06:27,471 --> 00:06:31,511
Doesn't that broad amount of input though
make it a challenge to come up with a

96
00:06:31,511 --> 00:06:34,823
consensus about what direction to go, what
the principle should be?

97
00:06:35,470 --> 00:06:36,830
It does.

98
00:06:36,850 --> 00:06:40,890
And I think that's part of the reasons
that we want to do that, because it is

99
00:06:40,890 --> 00:06:43,830
really important to keep all of the voices
included.

100
00:06:43,830 --> 00:06:48,990
And when you do that, and if you really go
for the diversity, of course we're going

101
00:06:48,990 --> 00:06:56,000
to hear inputs both sides of the spectrum
and trying to figure out the right space.

102
00:06:56,000 --> 00:06:58,970
And I don't want to say middle ground,
because there's a lot of the things to

103
00:06:58,970 --> 00:06:59,750
consider.

104
00:06:59,750 --> 00:07:04,590
You want to be scientific, high scientific
integrity and scientific validity, right?

105
00:07:04,590 --> 00:07:08,630
So we want to make sure that whatever we
put there is technically robust and

106
00:07:08,630 --> 00:07:09,950
empirically validated.

107
00:07:09,950 --> 00:07:16,010
It is operationalizable and implementable
today, but also it doesn't get expired in

108
00:07:16,010 --> 00:07:19,990
a very short amount of time.

109
00:07:20,030 --> 00:07:27,110
So we heard when we start this process
from everything and everywhere, from the

110
00:07:27,110 --> 00:07:31,350
civil society groups that wanted
everything to be absolutely

111
00:07:31,662 --> 00:07:38,842
public and open and audited and a lot of
vigorous tests and transparency

112
00:07:38,842 --> 00:07:41,402
requirements, a lot of really good reasons
for that.

113
00:07:41,402 --> 00:07:45,122
And we also hear from the technology
developers, importance of the IP and

114
00:07:45,122 --> 00:07:48,142
innovations and how those things have to
be protected.

115
00:07:49,222 --> 00:07:55,442
But that's the beauty of this, when you
give everybody the voice and we are

116
00:07:55,442 --> 00:08:00,622
objective and want to convene.

117
00:08:00,622 --> 00:08:08,382
try to make value for every participant in
the discussions to see how those consensus

118
00:08:08,382 --> 00:08:09,882
are being built.

119
00:08:09,882 --> 00:08:15,102
And the other thing is that when we get
input from a very vast group of the

120
00:08:15,102 --> 00:08:18,462
community, NIST has the pen in writing
this.

121
00:08:18,462 --> 00:08:21,102
So we would kind of make a decision where
we have to go.

122
00:08:21,102 --> 00:08:27,142
Of course, we go out and explain and make
sure that to the extent possible, we have

123
00:08:27,142 --> 00:08:27,722
the buy -in.

124
00:08:27,722 --> 00:08:29,646
Those buy -in are important because that's

125
00:08:29,646 --> 00:08:34,746
That's where the adoptions happens, but
that is part of the process, building the

126
00:08:34,746 --> 00:08:35,724
consensus.

127
00:08:36,879 --> 00:08:43,015
How can we measure the trustworthiness or
responsibility of an AI system?

128
00:08:44,334 --> 00:08:47,464
That's a very difficult, you know, that's
a technical challenge.

129
00:08:47,464 --> 00:08:50,374
That's a question that the community is
looking at this.

130
00:08:50,374 --> 00:08:54,014
It's difficult at many different
dimensions.

131
00:08:54,534 --> 00:08:58,914
So our institutions, our work, the group
that I have been a part of this, and by

132
00:08:58,914 --> 00:09:02,154
the way, I joined this in June 7, 1999.

133
00:09:02,154 --> 00:09:06,854
So in a few months, it's going to be my 25
years and all my career here, I've been

134
00:09:06,854 --> 00:09:13,102
working on computer vision, machine
learning type of efforts.

135
00:09:13,102 --> 00:09:16,062
and a lot of tests and evaluations.

136
00:09:16,062 --> 00:09:20,882
So if some of your listeners are familiar
with the face recognition tests or

137
00:09:20,882 --> 00:09:25,022
fingerprint work that we have done, I've
been part of that team doing that type of

138
00:09:25,022 --> 00:09:26,022
evaluation.

139
00:09:26,282 --> 00:09:32,342
So with a rich history and experience and
understanding of how to do evaluations,

140
00:09:33,222 --> 00:09:38,362
evaluations of the AI system poses a new
set of questions and challenges.

141
00:09:38,562 --> 00:09:41,646
One is, of course, what we mean by

142
00:09:41,646 --> 00:09:44,606
safe AI, what we mean by secure AI.

143
00:09:44,686 --> 00:09:52,486
But even if we can come up with some
understanding of those, AI systems are not

144
00:09:52,486 --> 00:09:57,226
just about data computing algorithm.

145
00:09:57,226 --> 00:10:01,186
There are complex interactions of the data
computing algorithm with the environment

146
00:10:01,186 --> 00:10:04,006
they're operating in, the human that
operates that.

147
00:10:04,006 --> 00:10:08,334
So the solutions and the answers are not
only technical and technology.

148
00:10:08,334 --> 00:10:12,054
but they all had to be answered through
the socio -technical lens.

149
00:10:12,054 --> 00:10:17,534
What that means is that the evaluations
that we have done before to get systems

150
00:10:17,534 --> 00:10:22,294
out of their context of use, bring them to
the laboratory, run the test data through

151
00:10:22,294 --> 00:10:24,644
them and test them, and those are
important.

152
00:10:24,644 --> 00:10:25,464
We have to do them.

153
00:10:25,464 --> 00:10:28,874
They give us some baseline of the
capabilities and limitations and failure

154
00:10:28,874 --> 00:10:34,134
modes, but they are not going to be
sufficient to surface all the sorts of the

155
00:10:34,134 --> 00:10:36,750
problems, negative consequences that can
come.

156
00:10:36,750 --> 00:10:38,410
out of the AI systems.

157
00:10:38,410 --> 00:10:43,370
We are hearing a lot about the red teaming
of AI systems with spot checking and

158
00:10:43,370 --> 00:10:50,510
basically trying to see how AI systems can
get into undesirable behavior.

159
00:10:51,330 --> 00:10:57,550
But beyond those things, we believe that
another set of testing where AI systems

160
00:10:57,550 --> 00:11:04,370
are being tested in their native context
of use with the humans, individuals that

161
00:11:04,370 --> 00:11:04,878
are

162
00:11:04,878 --> 00:11:06,518
going to use it, it's really important.

163
00:11:06,518 --> 00:11:07,598
We call them field testing.

164
00:11:07,598 --> 00:11:11,058
We have seen examples of that in sort of,
you know, pharmaceutical.

165
00:11:11,058 --> 00:11:12,838
FDA is doing this.

166
00:11:12,838 --> 00:11:20,318
So it's, there are all these difficulties
of exactly what we are testing for.

167
00:11:20,318 --> 00:11:24,738
And then the spectrum of the test that we
are doing is not going to be enough.

168
00:11:25,138 --> 00:11:27,088
The laboratory testing is not going to be
enough.

169
00:11:27,088 --> 00:11:29,598
We have to augment it with other testing.

170
00:11:29,598 --> 00:11:32,454
And then the...

171
00:11:32,622 --> 00:11:36,282
The other thing is that we are not going
to be able to test multi -purpose AI

172
00:11:36,282 --> 00:11:40,801
systems for all of the different use cases
that are being used.

173
00:11:40,801 --> 00:11:42,602
So how to scale all of these things?

174
00:11:42,602 --> 00:11:44,122
So a difficult problem.

175
00:11:44,122 --> 00:11:48,392
It is something that we have been working
on since the release of the AI RMF.

176
00:11:48,392 --> 00:11:52,362
One of the four functions on AI RMF is the
measure function.

177
00:11:52,362 --> 00:11:58,122
AI RMF provides, in addition to giving the
structured approach to what we mean by

178
00:11:58,122 --> 00:12:00,174
trustworthiness, the meat of that is...

179
00:12:00,174 --> 00:12:06,514
72 recommendations categorized in four
functions of map, measure, manage, and

180
00:12:06,514 --> 00:12:12,534
govern to how understand and identify AI
risks, how to measure them based on those

181
00:12:12,534 --> 00:12:17,494
measurement results, what type of risk
responses and mitigations are to be

182
00:12:17,494 --> 00:12:23,894
implemented, and the govern functions give
recommendations basically for the

183
00:12:23,894 --> 00:12:28,110
processes and procedures, all of the
embeddings that's needed for the right.

184
00:12:28,110 --> 00:12:30,930
risk management processes.

185
00:12:32,069 --> 00:12:36,970
So after the release of the AIRM, we put a
lot of our effort into the measure

186
00:12:36,970 --> 00:12:44,250
functions, which quite aligns and become
the foundation work that our latest

187
00:12:44,250 --> 00:12:47,770
assignment as part of the EO is built on
top of that.

188
00:12:47,770 --> 00:12:52,750
So a lot of work is happening as part of
the task that we have with the EO as part

189
00:12:52,750 --> 00:12:55,690
of the AI Safety Institute.

190
00:12:55,690 --> 00:12:57,006
We are standing up.

191
00:12:57,006 --> 00:13:01,366
and its consortium to work with the
community on figuring out the evaluations.

192
00:13:01,366 --> 00:13:06,186
The last thing I will say in this space is
that it's a tough problem and we are not

193
00:13:06,186 --> 00:13:08,426
going to get the answers in the first go.

194
00:13:08,426 --> 00:13:13,746
So we have to be on the understand that
it's going to be iterative and as we do

195
00:13:13,746 --> 00:13:17,986
more of this research, run evaluations,
you know, write guidelines based on those

196
00:13:17,986 --> 00:13:19,866
guidelines, conduct evaluations.

197
00:13:19,866 --> 00:13:24,266
The evaluations will tell us that what was
the gaps in the guidelines, the evaluation

198
00:13:24,266 --> 00:13:26,030
can also say the gaps in the

199
00:13:26,030 --> 00:13:31,410
So hopefully we are going to generate a
positive feedback loop that the

200
00:13:31,410 --> 00:13:38,330
measurement science is getting
strengthened by empirical evidence and

201
00:13:38,330 --> 00:13:40,450
science is going to be built around that.

202
00:13:40,450 --> 00:13:45,990
But also the practice and adoption of the
AIs are going to get stronger because of

203
00:13:45,990 --> 00:13:47,308
the test results.

204
00:13:47,824 --> 00:13:51,444
Can you talk about some of the directions
that you and the community are going to

205
00:13:51,444 --> 00:13:55,399
try to come up with those benchmarks for
some of those issues?

206
00:13:56,142 --> 00:14:03,542
All right, so October 30th, Executive
Order 14110, Safe, Secure, and Trustworthy

207
00:14:03,542 --> 00:14:06,022
AI was released.

208
00:14:06,022 --> 00:14:11,102
It's a very comprehensive set of
instructions to many of the federal

209
00:14:11,102 --> 00:14:16,262
government agencies on how to promote
safe, secure, trustworthy AI.

210
00:14:16,282 --> 00:14:23,790
Section 4 on safety and security tasks
several agencies, but predominantly NIST.

211
00:14:23,790 --> 00:14:30,170
on advancing the measurement for and
evaluations for safe, secure, trustworthy

212
00:14:30,170 --> 00:14:31,490
AI systems.

213
00:14:31,490 --> 00:14:42,260
In particular, it directs NIS to develop
guidelines for evaluations, reteaming,

214
00:14:42,260 --> 00:14:44,350
safety, cybersecurity.

215
00:14:44,350 --> 00:14:49,130
It asks NIS to facilitate development of
consensus -based standards.

216
00:14:49,130 --> 00:14:52,654
It also asks NIS to provide testing
environment for

217
00:14:52,654 --> 00:14:54,874
evaluations of the AI systems.

218
00:14:54,874 --> 00:14:58,814
These guidelines and infrastructures will
be voluntary resources for the community

219
00:14:58,814 --> 00:14:59,754
to use.

220
00:14:59,754 --> 00:15:04,534
So right now we had our heads down
drafting the documents, which we are

221
00:15:04,534 --> 00:15:07,134
hoping to go out for public comment.

222
00:15:07,134 --> 00:15:10,874
Based on the public comment, we're hoping
to improve it to meet the deadlines of the

223
00:15:10,874 --> 00:15:11,954
EO.

224
00:15:12,394 --> 00:15:19,214
And the EO gives us a 270 days deadline
that gets us to the end of the July.

225
00:15:19,214 --> 00:15:20,878
But all of these could work.

226
00:15:20,878 --> 00:15:25,358
All of these actions that the EEO sets in
the motion, we think that should not and

227
00:15:25,358 --> 00:15:27,798
cannot stop in end of July.

228
00:15:27,798 --> 00:15:34,978
So we're looking into the AI Safety
Institute and its consortium to give a

229
00:15:34,978 --> 00:15:38,898
long lasting approach to all of these
efforts in the research and evaluation

230
00:15:38,898 --> 00:15:39,878
builder.

231
00:15:40,118 --> 00:15:44,898
So what we are working with the community
is first develop the guidelines that we

232
00:15:44,898 --> 00:15:48,618
have asked through the EEO, working
through the consortium with the consortium

233
00:15:48,618 --> 00:15:49,390
members.

234
00:15:49,390 --> 00:15:56,650
when it's out to refine them, improve
them, conduct evaluations, and based on

235
00:15:56,650 --> 00:16:00,330
those kind of lived experiences, improve
the guidelines.

236
00:16:00,610 --> 00:16:06,070
Eventually, we definitely look like to
have sort of standardized test

237
00:16:06,070 --> 00:16:11,670
environments, standardized methods and
metrics and methodology for testing so

238
00:16:11,670 --> 00:16:16,690
that the community can use them and help
with the scaling of the evaluations that

239
00:16:16,690 --> 00:16:18,580
needed to be done in this space.

240
00:16:19,695 --> 00:16:25,055
Have you seen good examples of either
private sector companies or public sector

241
00:16:25,055 --> 00:16:28,743
organizations taking the RMF and
operationalizing it?

242
00:16:29,806 --> 00:16:33,606
We have heard a lot and we are really
grateful for that.

243
00:16:33,606 --> 00:16:40,726
So when the AIRMF was released January 26,
2023, a lot of different companies

244
00:16:40,726 --> 00:16:45,406
voluntarily and entities, I should say,
gave us some sort of testimonies on how

245
00:16:45,406 --> 00:16:48,456
they are planning to use AIRMF.

246
00:16:48,456 --> 00:16:52,146
And that again goes back to building that
engagement because when it comes out,

247
00:16:52,146 --> 00:16:54,446
everybody sees themselves in it.

248
00:16:54,806 --> 00:16:57,742
And then since then, we have heard from...

249
00:16:57,742 --> 00:17:02,162
large medium, small size businesses that
are using it.

250
00:17:02,682 --> 00:17:07,322
Some of those have given us, we call them
use cases, how they're using it

251
00:17:07,322 --> 00:17:07,752
internally.

252
00:17:07,752 --> 00:17:10,222
Some of those are posted on our website.

253
00:17:10,462 --> 00:17:14,942
Some of the other one has been talking to
us and letting us know that, particularly

254
00:17:14,942 --> 00:17:19,102
for the larger organizations, how they are
aligning their internal processes with AI

255
00:17:19,102 --> 00:17:19,862
or MF.

256
00:17:19,862 --> 00:17:24,270
And we're always happy to work with the
smaller size businesses to...

257
00:17:24,270 --> 00:17:26,650
help them with the operationalization.

258
00:17:26,650 --> 00:17:32,690
So one of the things, for example, we did
in March of 2023, we put out the

259
00:17:32,690 --> 00:17:34,610
trustworthy AI resource center.

260
00:17:34,610 --> 00:17:40,470
So in addition to the release of the
ARRMF, with the goal of keeping ARRMF, we

261
00:17:40,470 --> 00:17:46,450
call it as evergreen as possible, with a
revision cycle of three to five years, we

262
00:17:46,450 --> 00:17:51,110
put out a playbook out that basically
gives a lot of...

263
00:17:51,726 --> 00:17:56,866
informative references and information and
descriptions on exactly what to do for

264
00:17:56,866 --> 00:18:02,726
each of those 72 recommendations, what
they can do today, what are the tools and

265
00:18:02,726 --> 00:18:06,946
standards that can be used today for
implementations.

266
00:18:07,426 --> 00:18:12,206
So the playbook, we put a revision cycle
of every six months so we can keep it up

267
00:18:12,206 --> 00:18:19,086
with the latest changes in the technology
as tools and information becomes

268
00:18:19,086 --> 00:18:20,286
available.

269
00:18:20,966 --> 00:18:21,774
We...

270
00:18:21,774 --> 00:18:26,674
We made sure that this is sort of
interactive, filterable based on the role

271
00:18:26,674 --> 00:18:31,514
of the AI actor, designer, developer,
deployer, the area that they want to get

272
00:18:31,514 --> 00:18:35,714
information out from bias or privacy.

273
00:18:35,714 --> 00:18:44,114
And we're also working with the community
to add more tools and support, supportive

274
00:18:44,114 --> 00:18:48,814
documents, informative references,
implementations that can help with the

275
00:18:48,814 --> 00:18:49,694
implementation of that.

276
00:18:49,694 --> 00:18:51,470
But across the board, we have seen...

277
00:18:51,470 --> 00:18:52,750
It's being used.

278
00:18:52,750 --> 00:19:00,970
Another interesting or welcoming happening
is that at the time of the release of the

279
00:19:00,970 --> 00:19:04,050
AI RMF, we also released Crosswalks.

280
00:19:04,050 --> 00:19:08,690
So AI RMF was not the first document being
out, and we want to definitely make sure

281
00:19:08,690 --> 00:19:13,330
that we build based on the knowledge of
the community and leverage.

282
00:19:13,330 --> 00:19:18,610
For example, the definition of AI system
in the AI RMF is based on the OECD

283
00:19:18,610 --> 00:19:19,770
definition.

284
00:19:20,230 --> 00:19:21,038
So if we...

285
00:19:21,038 --> 00:19:22,558
based on the input of the community.

286
00:19:22,558 --> 00:19:27,338
We improved OECD definition and then OECD
took some of our improvement, put it back

287
00:19:27,338 --> 00:19:30,618
into their definitions, definition of AI
systems.

288
00:19:32,118 --> 00:19:38,578
So we want to make sure that we are not
adding confusion to the market by yet put

289
00:19:38,578 --> 00:19:39,628
another document out.

290
00:19:39,628 --> 00:19:45,798
So we put out crosswalks that says that,
for example, how AI RMF relates to OECD AI

291
00:19:45,798 --> 00:19:49,742
recommendations, the draft of the EU AI
Act.

292
00:19:49,742 --> 00:19:54,502
few weeks ago, but at that time it was a
draft with some other documents that's

293
00:19:54,502 --> 00:20:00,442
coming out of the White House, but also
with ISO, the standards that were coming

294
00:20:00,442 --> 00:20:04,182
out of the ISO international standard
organizations.

295
00:20:04,302 --> 00:20:10,002
ISO Insights is the technical tag to ISO,
US technical tag to ISO.

296
00:20:10,002 --> 00:20:14,042
Basically pick that up and as ISO is
coming up with other standards, they

297
00:20:14,042 --> 00:20:17,042
develop crosswalks with AI RMF.

298
00:20:17,142 --> 00:20:19,182
And hopefully all of these things that...

299
00:20:19,182 --> 00:20:25,142
I think is an indication that AI -RMF are
being used, so that's why they need those

300
00:20:25,142 --> 00:20:31,862
crosswalks, but also hopefully bringing
more semantic interoperability in the

301
00:20:31,862 --> 00:20:32,742
short run.

302
00:20:34,479 --> 00:20:39,619
You mentioned the executive order and the
AI Safety Institute that was one of the

303
00:20:39,619 --> 00:20:43,479
things that was created out of that
process and you're now taking over as CTO

304
00:20:43,479 --> 00:20:45,839
of that USAI Safety Institute.

305
00:20:45,919 --> 00:20:49,959
How is that different from the work that
NIST has done before and what's on the

306
00:20:49,959 --> 00:20:51,847
agenda for the Institute?

307
00:20:52,846 --> 00:21:04,506
Yeah, we look at both the EO and the
Safety Institute as basically, that's the

308
00:21:04,506 --> 00:21:11,586
phrase I use, supercharging our efforts in
advancing trustworthy responsible AI.

309
00:21:11,726 --> 00:21:19,246
So I think in terms of the goal and vision
to advance the science, practice, and

310
00:21:19,246 --> 00:21:20,910
policy of AI systems,

311
00:21:20,910 --> 00:21:25,770
increasing adoptions of the AI is the work
that we are doing.

312
00:21:25,770 --> 00:21:34,010
And all of these things are giving us more
capacity and strengthens our engagements

313
00:21:34,010 --> 00:21:36,570
to do better and more.

314
00:21:36,590 --> 00:21:41,550
Eventually, we also want to keep an eye on
the, or help with the adoption.

315
00:21:41,550 --> 00:21:45,690
So as a measurement science agency, we
want to make sure that the right

316
00:21:45,690 --> 00:21:49,934
measurement science tools, standards,
guidelines are there.

317
00:21:49,934 --> 00:21:53,094
for anybody that wants to design, develop,
or use AI systems.

318
00:21:53,094 --> 00:21:59,534
And hopefully, by doing this, we actually
help the adoption rates of AI in science,

319
00:21:59,554 --> 00:22:06,774
in health, climate, agriculture, all of
those areas that we think that that's

320
00:22:06,774 --> 00:22:12,756
where AI technology can bring many, many
benefits.

321
00:22:13,967 --> 00:22:17,447
What do you think the biggest gap is right
now or the biggest area where we haven't

322
00:22:17,447 --> 00:22:21,647
seen enough progress on some of the
elements that are going to be necessary to

323
00:22:21,647 --> 00:22:22,567
promote that?

324
00:22:23,886 --> 00:22:28,866
I might be biased coming from a
measurement science agency, but I would

325
00:22:28,866 --> 00:22:30,946
say measurement and evaluations.

326
00:22:30,946 --> 00:22:34,606
I think everybody wants to save, secure,
trustworthy AI, but we don't know how to

327
00:22:34,606 --> 00:22:34,996
measure them.

328
00:22:34,996 --> 00:22:36,286
We don't have the metrics there.

329
00:22:36,286 --> 00:22:38,886
We don't have the methodologies there.

330
00:22:39,266 --> 00:22:48,706
There are documents, policies they like to
talk about, assuring that AI systems are

331
00:22:48,706 --> 00:22:48,946
safe.

332
00:22:48,946 --> 00:22:50,646
You know, we saw it in the EU AI Act.

333
00:22:50,646 --> 00:22:51,950
They want to see...

334
00:22:51,950 --> 00:22:56,570
some sort of assurance or conformity
assessment, but those specifications

335
00:22:56,570 --> 00:22:58,790
hasn't been written yet.

336
00:22:58,790 --> 00:23:02,090
So what do we mean by safety?

337
00:23:03,210 --> 00:23:09,110
And policies usually stay at the abstract
high level that AI should be safe or non

338
00:23:09,110 --> 00:23:14,270
-discriminatory, but bringing it down,
getting everybody on the same page, that's

339
00:23:14,270 --> 00:23:17,050
what it means, and how to assure for them.

340
00:23:17,050 --> 00:23:21,358
I think that's the biggest, that's one of
the biggest gaps.

341
00:23:22,670 --> 00:23:30,450
In general, in this AI space, I will say
that we know a lot less than what we

342
00:23:30,450 --> 00:23:31,290
should.

343
00:23:31,290 --> 00:23:39,190
So anything that we can do to be able to
know more what to expect, how to

344
00:23:39,190 --> 00:23:45,450
characterize the behavior functionality of
the systems, these are the good partners.

345
00:23:47,375 --> 00:23:51,375
How much are those really scientific
questions though?

346
00:23:51,375 --> 00:23:55,494
And part of what I'm thinking about is you
get to some issues like bias, where

347
00:23:55,494 --> 00:23:59,555
there's various different tests about
algorithmic fairness, but some of them are

348
00:23:59,555 --> 00:24:00,915
fundamentally incompatible.

349
00:24:00,915 --> 00:24:04,675
And you mentioned the socio -technical
aspects of the systems where there may be

350
00:24:04,675 --> 00:24:07,765
things that are not just about yes or no.

351
00:24:07,765 --> 00:24:14,335
So how do you and others in the community
work to identify what it is where we can

352
00:24:14,335 --> 00:24:17,575
actually get an answer through the kind of
science you're talking about?

353
00:24:17,710 --> 00:24:18,650
Yeah.

354
00:24:18,650 --> 00:24:20,350
Thank you for that question.

355
00:24:20,350 --> 00:24:29,230
So back in March 2022, we put out a
document out, special publication 1270.

356
00:24:29,230 --> 00:24:36,770
It has a name towards standards for
understanding harmful bias in AI systems.

357
00:24:36,770 --> 00:24:40,550
But these are some of the foundational
work that we're doing with the community

358
00:24:40,550 --> 00:24:41,382
that.

359
00:24:41,486 --> 00:24:45,706
Again, as you said, there are tools for
bias, but what are really measuring?

360
00:24:45,706 --> 00:24:46,856
Are they measuring the right thing?

361
00:24:46,856 --> 00:24:48,286
What should we really measure?

362
00:24:48,286 --> 00:24:51,846
What is bias across the life cycle of AI
anyway?

363
00:24:51,846 --> 00:24:54,496
So that document talks about a lot of
these things.

364
00:24:54,496 --> 00:24:58,546
And if I want to summarize, the most
important to me, the most important point

365
00:24:58,546 --> 00:25:07,806
of that document is that bias is not just
about disparity and demographic disparity.

366
00:25:07,806 --> 00:25:09,582
A lot of us, when you're talking about

367
00:25:09,582 --> 00:25:12,962
bias, the first thing that comes to our
mind is demographic disparity.

368
00:25:12,962 --> 00:25:19,422
The document talks about computational and
statistical bias as one type of bias, but

369
00:25:19,422 --> 00:25:22,802
there are also other types of bias,
systemic biases.

370
00:25:22,802 --> 00:25:28,982
If you go and do a statistically valid
data collection of the community and

371
00:25:28,982 --> 00:25:34,622
society today, that's going to be biased
and the biases baked into the society will

372
00:25:34,622 --> 00:25:39,214
get into the systems because our training
data will carry those biases.

373
00:25:39,214 --> 00:25:43,934
It also talks about another type of bias,
which is cognitive biases.

374
00:25:43,934 --> 00:25:49,534
How the different individuals process
information given by AI systems or

375
00:25:49,534 --> 00:25:51,374
interact with AI systems.

376
00:25:51,374 --> 00:25:54,394
And based on that, how we react to those
is different.

377
00:25:54,394 --> 00:25:56,234
So there's biases there too.

378
00:25:56,234 --> 00:26:02,254
Understanding all of these types of
biases, having a wholesome understanding

379
00:26:02,254 --> 00:26:05,678
of now in this topic what the bias is
and...

380
00:26:05,678 --> 00:26:11,738
how biases showed itself at the different
stages of the AI lifecycle become really

381
00:26:11,738 --> 00:26:16,158
important on other steps of the
evaluation.

382
00:26:16,158 --> 00:26:21,378
So the science part of that, again,
becomes to having a sort of a wholesome

383
00:26:21,378 --> 00:26:27,578
approach to what could be the sources of
the biases, where they're coming from, and

384
00:26:27,578 --> 00:26:32,174
that information can help us to do the
measurement.

385
00:26:32,174 --> 00:26:35,774
You asked the question about the socio
-technical.

386
00:26:35,774 --> 00:26:38,644
So all of these problems are socio
-technical.

387
00:26:38,644 --> 00:26:41,934
So there is not just a technical or
technology solution for them.

388
00:26:41,934 --> 00:26:46,474
We just talk about cognitive science
biases, for example.

389
00:26:46,934 --> 00:26:54,774
And the fact that one of the components of
the responsible AI is human centricity.

390
00:26:54,934 --> 00:26:57,354
So you have the humans there.

391
00:26:57,374 --> 00:27:01,742
And that gets to another challenge that we
have with building the science.

392
00:27:01,742 --> 00:27:08,042
for the measurement that a lot of things
that we are used to on, for example,

393
00:27:08,042 --> 00:27:15,522
measuring length, mass, electricity, those
areas that NIST work on coming up with

394
00:27:15,522 --> 00:27:21,642
standards for measuring them at the turn
of the last century, which were critical

395
00:27:21,642 --> 00:27:25,642
for the underlying technologies of that
time.

396
00:27:25,782 --> 00:27:27,374
Right now we have...

397
00:27:27,374 --> 00:27:33,174
quantitative metrics to measure them,
intropable ways of measuring them.

398
00:27:33,814 --> 00:27:38,794
We are at the distant of technology, we
are trying to do the very same thing for

399
00:27:38,794 --> 00:27:40,640
the technology of our time.

400
00:27:42,511 --> 00:27:45,931
As you mentioned, some other
jurisdictions, the European Union in

401
00:27:45,931 --> 00:27:49,775
particular, have now a comprehensive
legislation.

402
00:27:52,047 --> 00:27:53,087
states.

403
00:27:53,747 --> 00:27:59,147
Talk a little bit about what the function
is of the work that NIST does in that

404
00:27:59,147 --> 00:28:03,015
context where there aren't necessarily
specialized regulations.

405
00:28:03,694 --> 00:28:07,314
So NIST is a non -regulatory agency.

406
00:28:10,134 --> 00:28:13,464
We don't work with regulators.

407
00:28:13,464 --> 00:28:17,014
None of the things that we put out is for
mandatory use.

408
00:28:17,014 --> 00:28:18,694
Everything is for voluntary use.

409
00:28:18,694 --> 00:28:25,174
And even the evaluations that we do while
we evaluate for, for example, accuracy or

410
00:28:25,174 --> 00:28:28,430
security or privacy or any of these
things.

411
00:28:28,430 --> 00:28:33,130
We don't make any judgment on how safe is
safe, how private is private, how accurate

412
00:28:33,130 --> 00:28:37,830
is accurate, because there is no one size
fits all and it depends on the context of

413
00:28:37,830 --> 00:28:38,490
use.

414
00:28:38,490 --> 00:28:44,770
What we are trying to do and our work with
the policy makers is to provide empirical

415
00:28:44,770 --> 00:28:51,810
evidence to help with development of the
policies that are clear and also make sure

416
00:28:51,810 --> 00:28:57,134
that the underlying science and technical
tool and standard exists.

417
00:28:57,134 --> 00:29:00,894
to help with the enforcement of the
policy.

418
00:29:00,894 --> 00:29:05,594
We just talked about it, that the laws can
come and say that AI must be safe, but

419
00:29:05,594 --> 00:29:06,354
safety means.

420
00:29:06,354 --> 00:29:13,384
And we see it what happened with the EU
Act as well that after its passing, now,

421
00:29:13,384 --> 00:29:17,854
since then, like European centralization
body are tasked with coming up with the

422
00:29:17,854 --> 00:29:23,714
standards for basically checking the
conformity assessment to the requirements

423
00:29:23,714 --> 00:29:24,654
in the law.

424
00:29:24,654 --> 00:29:25,070
So...

425
00:29:25,070 --> 00:29:26,830
So we are not regulatory.

426
00:29:26,830 --> 00:29:29,630
Everything that we put out is for
voluntary use.

427
00:29:29,630 --> 00:29:37,330
But we give sort of empirical evidence for
clear policy making, evidence -based

428
00:29:37,330 --> 00:29:44,010
policy making, and the technical tools and
the measurements and evaluations needed

429
00:29:44,010 --> 00:29:48,230
and standards needed for enforcement of
that.

430
00:29:48,230 --> 00:29:51,470
I use the word standards several times,
and it may be.

431
00:29:51,470 --> 00:29:57,090
helpful if I clear that while standard is
our middle name, we are not standard

432
00:29:57,090 --> 00:30:00,670
development organizations such as ISO or
IEEE.

433
00:30:00,670 --> 00:30:06,370
There are really good policies in the US
in effect and in existence that the job of

434
00:30:06,370 --> 00:30:11,400
the federal government is to help industry
develop the standards.

435
00:30:11,400 --> 00:30:13,310
So they are in the driver's seat.

436
00:30:13,310 --> 00:30:19,530
We are supporting their work because,
again, we bring those objective technical

437
00:30:19,530 --> 00:30:20,558
contributions.

438
00:30:20,558 --> 00:30:22,718
in the development of the standards.

439
00:30:23,118 --> 00:30:29,298
Back in 2019, we had a task by an
executive order to write a plan for

440
00:30:29,298 --> 00:30:33,238
federal government engagement in
development of AI standards.

441
00:30:33,238 --> 00:30:40,438
With this current EL, we have a task to
develop a plan for global engagement to

442
00:30:40,438 --> 00:30:43,698
promote consensus industry standards.

443
00:30:44,578 --> 00:30:45,358
And...

444
00:30:45,358 --> 00:30:49,398
So we work a lot with the community, both
industry and standard development

445
00:30:49,398 --> 00:30:53,938
organizations on promotion and facilitate
development of the good standards, but we

446
00:30:53,938 --> 00:30:56,520
are not developing standards.

447
00:30:57,199 --> 00:31:02,939
What is your sense about the current state
of global cooperation and coordination on

448
00:31:02,939 --> 00:31:04,039
AI response?

449
00:31:05,390 --> 00:31:11,110
So a lot of really, a lot of conversations
has been happening for the past many years

450
00:31:11,110 --> 00:31:14,610
that I have been involved and a lot of
really good work is happening.

451
00:31:14,750 --> 00:31:21,990
OECD has, OECD AI Governance Party is
doing a lot of good work on bringing

452
00:31:21,990 --> 00:31:25,850
communities together but also provide a
lot of technical clarity for policy

453
00:31:25,850 --> 00:31:29,490
meeting through a lot of the work that
they're doing, a lot of documents that

454
00:31:29,490 --> 00:31:31,370
they're putting out.

455
00:31:31,490 --> 00:31:33,370
We saw.

456
00:31:33,582 --> 00:31:40,142
a really good progress through the G7 kind
of code of conduct that try to bring more

457
00:31:40,142 --> 00:31:49,542
sort of a global interoperability on sort
of voluntary commitment that happened in

458
00:31:49,542 --> 00:31:50,842
the US.

459
00:31:51,082 --> 00:31:56,402
Bottom line is that global conversations
and global interoperability for the AI

460
00:31:56,402 --> 00:31:58,682
space is really important.

461
00:31:58,682 --> 00:32:02,542
It is for any type of standards, but
particularly in the space.

462
00:32:02,542 --> 00:32:10,362
The technology is obviously very global
and the users are across the globe.

463
00:32:10,362 --> 00:32:17,102
So we are thinking about really maximizing
the benefits of AI while minimizing its

464
00:32:17,102 --> 00:32:21,562
negative consequences, harnessing the
power of AI for good.

465
00:32:21,622 --> 00:32:30,638
We need more global engagement, both at
the research side to build.

466
00:32:30,638 --> 00:32:34,138
the scientific underpinning on the
standard side.

467
00:32:34,138 --> 00:32:41,078
So we all are kind of using the same
standards and also on the implementation

468
00:32:41,078 --> 00:32:42,818
and operationalization side.

469
00:32:42,818 --> 00:32:46,338
So we can all benefit from this
technology.

470
00:32:46,338 --> 00:32:51,138
So a lot of really good work is happening
and a lot more can be done, I guess.

471
00:32:51,855 --> 00:32:52,285
Absolutely.

472
00:32:52,285 --> 00:32:54,435
Well, that's a really great place to wind
up.

473
00:32:54,435 --> 00:32:56,863
Thank you so much for a really fascinating
conversation.

474
00:32:57,262 --> 00:32:58,232
Thanks for having me.

475
00:32:58,232 --> 00:32:59,618
Enjoy the conversation.

