見出し画像

プロメテウスの火を継ぐ者たち——チューリングからChatGPTまで、AI八十六年の物語

プロメテウスの火を継ぐ者たち
——チューリングからChatGPTまで
AI八十六年の物語

目次

本書について
序章
第一章 半分の林檎
第二章 櫓も舵もなき船
第三章 蝋の翼
第四章 隠れた灯火
第五章 黒船とサムライ
第六章 夜明けの工房
第七章 二つの神の一手
第八章 八人の使徒
第九章 西へ
第十章 プロメテウスの火
あとがき
用語解説集
 

本書について

本書は、論文・公的記録・歴史資料を土台に、AIの発展を人物の物語として描く歴史読み物である。人物の動作や場面の空気には、史実から再構成した文学的描写を含む。一方、史料に逐語記録のない会話や内面は、確認された発言としてではなく、場面の意味を伝えるための再構成として扱う。

序章

人類は古くから、自らの似姿を作る夢を見てきた。
粘土から。
泥から。
象牙から。
そして、ときには死者の肉から。
世界中の神話と物語が、その夢を語っている。
ピグマリオンの彫った象牙の像は、女神アフロディーテの恵みによって息をした。プラハのラビは粘土からゴーレムを造り、その額に『真実』の語を刻んで命を与えた。十九世紀の物語では、スイスに生まれた若き自然哲学者が、墓場から集めた肉片を縫い合わせ、生きるものを作り上げた。
なぜ人類は、このような物語を繰り返し語ってきたのか。
おそらく私たちは、自分たちが『考える』ということ、『生きる』ということの不思議を、もう一人の自分を作る夢によって問い続けてきたのである。
命とは何か。
心とは何か。
言葉を発し、応答し、選び、誤り、悩むものを、私たちは作ることができるのか。
作られたものが目を開くとき、問われるのは、その怪物や人形の正体だけではない。作った私たち自身が、何者であるのかもまた、問い返される。
二十世紀になり、その夢は神話の炉を離れ、数学と工学の問題になった。
一九三六年、ケンブリッジ。
二十四歳を目前にした若き数学者アラン・チューリングは、紙と記号と一本のテープだけで、計算する機械の概念を描いた。そこにはまだ、金属の身体も、声も、顔もなかった。だが、その抽象的な機械の奥には、のちに『考える機械』と呼ばれるものの原型があった。
そこから、八十六年。
二〇二二年十一月三十日。
世界中の人々が、自分の手の中の小さな機械を通じて、言葉を返すAIと、文字通り対話を始めた。
これは、その八十六年の物語である。
それは、勝利だけの物語ではない。
闘いの物語であり、敗北の物語であり、長い冬ののちに訪れる復活の物語である。
ある者は、半分の林檎を残して去った。
ある者は、蝋の翼で空へ挑み、海に落ちた。
ある者は、研究室の灯火を、冬の中で消さずに守り続けた。
ある王者は、盤上で機械に敗れた。
ある棋士は、機械の読みを越える一手を放った。
ある八人は、一篇の論文を残し、世界へ散っていった。
そしてある日、誰かが神々の火を、ひそかに人類へ渡した。
登場する者たちの誰一人として、自分たちが何を作っているのかを、その時点で完全には理解していなかった。
それでも彼らは、進んだ。
一歩、また一歩。
暗い海へ。
見えない空へ。
まだ地図のない岸へ。
そしていま、私たちはその岸辺に立っている。
向こうに何が広がっているのか——まだ、誰も知らない。
これは、その物語である。

第一章 半分の林檎

アラン・チューリングの肖像(一九一二—一九五四)

一 ケンブリッジの自転車

一九三六年五月、ケンブリッジ。
キングス・カレッジの裏手から伸びる細い道を、ひとりの若い男が自転車で走っていた。背は高く、痩せている。風に乱れた髪。胸ポケットには万年筆。ハンドルを握る指は細く長く、どこか神経質に見えた。
アラン・マシスン・チューリング。二十四歳を目前にした若き数学者である。前年、彼はわずか二十二歳でキングス・カレッジのフェローに選ばれていた。誰もがその才能を認めていたが、その才能がどこへ向かおうとしているのか、まだ誰にも見えてはいなかった。
春の終わりの花弁が芝に散っていた。だが彼は、それをほとんど見ていなかった。頭の中では、まだこの世に存在しない機械が動いていた。
長い紙テープ。
その上を這う、ひとつの読み書き装置。
記号を読み、状態を変え、記号を書き、右か左へ一目盛り動く。
ただ、それだけ。
けれど、その『ただそれだけ』の中に、彼は計算という営みの骨格を見ていた。人間が紙と鉛筆を前にして行う計算を、極限まで削ぎ落とせば何が残るのか。記号を読むこと。規則に従うこと。次の状態へ移ること。ならば、計算とは天才のひらめきではなく、手続きとして記述できるものではないか。
彼はその抽象的な機械を、『自動機械』と呼ぶことにした。
論文の題は、すでに決まっていた。
『計算可能な数について、決定問題への応用』
それは、当時の多くの人々にとっては、数学の奥まった部屋で交わされる難解な議論にすぎなかった。だが、のちの時代は知ることになる。そこに描かれていたのは、現代のコンピュータの原型であり、さらに遠く、人工知能という夢の最初の輪郭であった。
ガリレオが望遠鏡で天を覗き、地球が宇宙の中心ではないことを見たように、チューリングは計算する者の手元を覗き込み、知性の一部が手続きへと還元できることを見た。
その発見は、まだ誰をも脅かしてはいなかった。
まだ、異端ではなかった。
このときの彼には、自分の頭の中の小さな機械が、やがて世界中の机の上に、工場に、病院に、兵器に、そして人々の掌の中に棲みつくことなど、想像できなかっただろう。
彼はただ、美しい絵を見ていた。
一本のテープと、ひとつの頭部。
それだけで世界を書き換える、静かな絵を。

二 ブレッチリーの小屋

三年後、世界は戦争の中にあった。
バッキンガムシャー、ブレッチリー・パーク。古い屋敷の周囲には、急ごしらえの木造小屋がいくつも建てられていた。優雅な芝生と、湿った板壁。その奇妙な取り合わせの中で、英国中から集められた数学者、言語学者、チェス愛好家、クロスワード・パズルの名人たちが、鉛筆と紙と紅茶を武器に戦っていた。
彼らが相手にしていたのは、ドイツ軍の暗号である。
チューリングが通ったのは、ハット八番。海軍暗号を扱う部署だった。
北大西洋では、ナチス・ドイツの潜水艦が英国へ向かう商船を次々に沈めていた。食料も燃料も兵器も、海を渡らなければ英国には届かない。その海を支配するため、ドイツ軍は『エニグマ』と呼ばれる暗号機を用いていた。
エニグマは、美しい機械だった。複数の回転盤が電気回路を組み替え、設定は日ごとに変わる。可能な組み合わせは天文学的な数に達した。人間の手で総当たりするなど、事実上不可能である。
だがチューリングは、その『不可能』を論理の問題として見つめた。
ポーランドの数学者たちが切り開いた突破口を受け継ぎながら、彼は同僚たちとともに、エニグマに対抗する機械を構想した。のちに『ボンベ』と呼ばれる装置である。
それは答えを考える機械ではない。
人間が見つけた論理の隙間を、疲れを知らずに辿る機械だった。
電動の歯車が回り、可能な設定を凄まじい速度で試していく。一日の試行は、人間の何百年分にも相当した。やがて、解けないはずの暗号は解けはじめる。
ドイツ軍は知らなかった。自分たちの最高機密の通信が、英国の片田舎の小屋で読まれていることを。大西洋の海の下で交わされた命令が、木造小屋の机の上で、紙片となって広げられていることを。
戦後の研究では、ブレッチリー・パークの暗号解読が戦争を数年短縮し、無数の命を救ったとも言われる。だが、その事実は長く国家機密として封じられた。
世界は、チューリングが何をしたのかを知らなかった。
世界は、自分が彼に救われたことを、彼の生きている間には知らなかった。
同僚たちは、彼の奇妙な習慣を覚えている。マグカップを盗まれないよう、ラジエーターに鎖でつないでいたこと。花粉症がひどく、自転車に乗るとき防毒マスクをかぶっていたこと。長距離を走らせれば、オリンピック選手に迫るほど速かったこと。
奇妙な男だった、と人々は言った。
だが同時に、誰もが知っていた。
あの奇妙な男がいなければ、戦争はもっと長く続いていたかもしれない、と。

三 模倣ゲーム

戦争が終わって五年が過ぎた。一九五〇年。
哲学誌『マインド』に、一篇の論文が掲載された。題は簡潔だった。
『計算機械と知能』
冒頭で、チューリングは問う。
『機械は考えることができるか』
だが彼はすぐに、その問いの形を変えた。『考える』とは何かを定義しようとすれば、議論は果てしなく続く。心とは何か。意識とは何か。魂とは何か。そうした問いは、人類が千年以上も抱えてきた迷宮だった。
チューリングは、その迷宮へ正面から入らなかった。代わりに、ひとつのゲームを提示した。
別々の部屋に、人間と機械がいる。判定者は文章だけで両者と会話する。声も顔も見えない。文字だけが行き交う。判定者が、人間と機械を区別できないならば——その機械は、考えているとみなしてよいのではないか。
後に『チューリング・テスト』と呼ばれる思考実験である。
彼の鋭さは、答えを出したことではない。問いの置き場所を変えたことにあった。知能を神秘として崇めるのではなく、観察できる振る舞いとして扱う。内側に魂があるかどうかではなく、外側に現れる応答を問題にする。
それは、哲学の問いを工学の手に渡す行為だった。
論文の中で、彼は予言した。五十年ほどのうちに、人々は『機械が考える』と語ることに、さほど抵抗を覚えなくなるだろう、と。
その予言が本当に日常の風景となるまでには、さらに時間がかかった。だが方向は間違ってはいなかった。七十年ほど後、人々は小さな画面越しに、機械と文章で会話するようになる。
相談し、質問し、冗談を言い、怒り、励まされるようにさえなる。
チューリングは、その光景を見ることがなかった。

四 二月の手紙

一九五二年一月の終わり、マンチェスター。
チューリングの自宅に、ある夜、若い男が滞在した。名をアーノルド・マレーという。数日後、家に泥棒が入った。チューリングは警察に通報した。
警察は盗難を調べるうちに、別の事実へたどり着いた。
当時の英国では、男性同士の関係は『重大わいせつ罪』として処罰の対象だった。チューリングは、自分の関係を隠そうとはしなかった。彼にとって、それは隠すべき悪事ではなかったからである。
しかし法は、そうは見なかった。
戦時中の仕事の多くは、なお国家機密の奥にあった。彼がブレッチリー・パークで果たした役割は、公の法廷で彼を守る勲章にはならなかった。国家に尽くした功績と、国家が定めた法とは、別々の秤に載せられた。
裁判が始まり、彼は有罪となった。
刑は、投獄か、あるいは当時『治療』と呼ばれたホルモン投与か。彼は後者を選んだ。合成女性ホルモンの投与は、彼の身体を変えていった。胸が膨らみ、筋力が落ちた。走ることを愛していた身体は、以前のようには動かなくなった。
それでも、仕事は続いた。
裁判を目前にした二月、友人ノーマン・ラウトリッジに宛てた手紙の中で、彼はひとつの三段論法を書いた。
チューリングは、機械が考えると信じている。
チューリングは、男と寝る。
ゆえに、機械は考えない。
痛烈な諧謔だった。
論理学者である彼は、世間がいかに粗雑な論理で人を裁くかを知っていた。そしてその粗雑さを、もう一度、論理の形に整えて笑い返したのである。
三百年余り前、ガリレオ・ガリレイは宗教裁判にかけられ、地動説を撤回させられた。彼が裁かれたのは、星を動かしたからではない。見たものを、見たとおりに語ったからだった。
チューリングが裁かれたのは、機械についての思想ではない。だが、権威が現実よりも規範を優先するとき、人は見えているものを見えていると言うだけで罪にされる。
伝説によれば、ガリレオは法廷を出たのち、小さく『それでも、地球は動く』と呟いたという。
チューリングが何を呟いたのか、史料は何も伝えていない。
ただ、彼が書いた論文は残っている。
そこでは機械が、静かに考えはじめていた。

五 半分の林檎

一九五四年六月七日の夜。
チェシャー州ワイムスロウ。チューリングの自宅。
翌朝、家政婦が彼の遺体を発見した。四十一歳だった。
そばには、半分かじられた林檎があった。
検査の結果、彼の体内からはシアン化合物が検出された。公式には自死とされた。ただし、林檎そのものにシアン化合物の検査がなされた記録はない。彼は自宅で化学実験を行っており、シアン化合物を扱う可能性もあった。事故であった可能性を指摘する声も、後に現れる。
真相は、どこにあるのか。
それでも、半分の林檎は残った。
それは証拠品である前に、ほとんど伝説のような姿を帯びていた。チューリングはディズニーの『白雪姫』を好み、王妃が毒林檎を作る場面の歌を口ずさんでいたとも伝えられる。
世界を救った男を、その世界が殺した。
そう書けば、物語はわかりやすい。
しかし歴史は、わかりやすい物語だけでできてはいない。彼の死には、事故の可能性も、孤独の影も、当時の制度の暴力も、いくつもの層が重なっている。
ただ、ひとつだけ確かなことがある。
世界は、彼に十分な敬意を払わなかった。
彼の死は、当時の新聞では小さく報じられただけだった。数学者が自宅で死亡した。それだけである。彼が何を成し遂げたのか。どれほど多くの命が彼の仕事に救われたのか。彼の一九三六年の抽象機械が、やがて世界そのものを変えることになるのか。
世界は、まだ知らなかった。
半分の林檎だけが、枕元に残されていた。

六 半世紀ののち

二〇〇九年。
英国首相ゴードン・ブラウンは、政府を代表してチューリングへの公式謝罪を発表した。彼は、もっとよい扱いを受けるべきだった、と。
二〇一三年、エリザベス女王はチューリングに特赦を授与した。
二〇一七年、彼と同じく過去の同性愛行為によって有罪とされた数万人の男性たちに、遡及的な恩赦が与えられた。世にいう『チューリング法』である。
二〇二一年には、英国の五十ポンド紙幣に彼の肖像が刻まれた。
かつて罪人とされた男の顔が、国家の通貨となった。
ガリレオが教皇庁から事実上の名誉回復を受けるまでには、三百五十九年を要した。チューリングは、半世紀余りで名誉を回復した。
半世紀で済んだ、と言うべきだろうか。
それとも、半世紀もかかった、と言うべきだろうか。
歴史は自然に正しくなるわけではない。誰かが記憶し、誰かが問い直し、誰かが謝罪を求める。その積み重ねの果てに、ようやく過去は少しだけ向きを変える。
それでも、向きは変わった。
かつて彼を裁いた国は、やがて彼の顔を紙幣に刻んだ。
かつて彼の名を隠した世界は、やがて彼を計算機科学の父と呼んだ。

終 残された絵

チューリングが遺したものは、ひとつの機械ではない。
ひとつの絵である。
長いテープ。記号。状態。規則。読み、書き、移動する頭部。
それはあまりに単純で、ほとんど貧しいほどだった。だが、その貧しさの中に、現代の計算機のすべてが宿っていた。コンピュータも、暗号も、プログラムも、そして人工知能も、この絵の遠い変奏である。ChatGPTも、Claudeも、Geminiも、AlphaGoも、もはや紙テープの上を動いているわけではない。シリコンの回路を走り、巨大なデータの海を渡り、確率の網の中から言葉や手を選び取っている。
それでも、その根底にはまだ、チューリングの絵がある。
規則に従って状態を変えるもの。
記号を読み、記号を書き、次の一手へ進むもの。
人間が『考える』と呼んできた振る舞いを、機械の側から照らし返すもの。
彼は、半分の林檎を残して去った。
残された半分とは、死の謎だけではない。
人間の知性とは何か。機械はどこまで考えうるのか。私たちは、自分たちだけのものだと思っていた思考を、どこまで機械と分かち合うのか。
その問いの半分が、いまも残っている。
毒であったのか、果実であったのか。
喪失であったのか、始まりであったのか。
チューリングの半分の林檎は、二十世紀の底へ沈み、二十一世紀の掌の上に浮かび上がった。
それでも、機械は考えるのか。
彼の残した絵は、いまも静かに動いている。
第一章 了

第二章 櫓も舵もなき船

ダートマス会議の夏(一九五六)

一 ハノーバーの提案書

一九五五年八月、ハノーバー。
ニューハンプシャー州の片隅にある、人口数千の小さな町。ダートマス大学数学科の若き助教授ジョン・マッカーシーは、二十八歳になろうとしていた。
この年の初め、カリフォルニアのスタンフォード大学から、この東海岸の田舎大学へ移ってきたばかりである。
その夏、彼は一つの名前を紙の上に置いた。
Artificial Intelligence。
人工知能。
一九五五年八月三十一日付でまとめられた研究計画書の冒頭には、マッカーシー、マービン・ミンスキー、ナサニエル・ロチェスター、クロード・シャノンの四人の名が並んでいた。そして、まだ学問分野として定着していなかったその言葉が、はっきりと記されていた。
『人工知能』。
似た言葉なら、すでにあった。サイバネティクス。自動制御。情報理論。複雑情報処理。
だが、マッカーシーが欲しかったのは、もっと広い旗だった。
機械に計算をさせるだけではない。
機械に、知的なふるまいそのものをさせる。
学習も、言語も、抽象化も、問題解決も、その下に集めることのできる名前。
研究計画書は、夏の二か月、十人ほどの研究者をダートマス大学に集める構想を掲げた。知能のあらゆる側面は、原理的には精密に記述でき、その記述をもとに機械で模倣できるはずだ——という、驚くほど大胆な仮説が、その中心にあった。
二か月。十人。知能の解明。
いま振り返れば、ほとんど無謀なほどの楽観である。だが、始まりというものは、たいてい楽観を燃料にしている。まだ海図がない者だけが、遠い海へ漕ぎ出せることもある。
計画は、当初の構想どおりには進まなかった。ロックフェラー財団から認められた助成は、五週間分、七千五百ドルだった。それでも十分だった。
歴史は、しばしば、こうした小さな数字から始まる。
一枚の提案書。
夏の数週間。
十人ほどの研究者。
七千五百ドル。
そして、まだ中身の定まっていないひとつの名前。
人工知能。

二 夏の到着

一九五六年の夏、ダートマス大学。
研究者たちは、ハノーバーの町へ少しずつ集まってきた。全員が同じ日に顔をそろえたわけではない。そもそも、それは今日われわれが想像するような、整然とした国際会議ではなかった。議題が細かく決められ、発表時間が割り振られ、議事録が整えられるような場ではない。
それはむしろ、若い学問のための寄合だった。
マービン・ミンスキー。二十八歳。脳と機械の境界に取り憑かれた若者。
クロード・シャノン。四十歳。情報という見えないものに数学の骨格を与えた巨人。
ナサニエル・ロチェスター。IBMの研究者であり、最新鋭計算機の設計に関わった実務家。
アーサー・サミュエル。チェッカーを指すプログラムを作りながら、機械が経験から強くなる可能性を探っていた。
レイ・ソロモノフ、オリバー・セルフリッジ、トレンチャード・モア。
そして、ハーバート・サイモンとアレン・ニューウェル。
参加者は時期をずらして出入りし、議論は部屋の中だけに収まらなかった。食堂へ、芝生へ、夜の町へと流れていったことだろう。
彼らが語っていたものは大きかった。
機械翻訳。チェス。定理証明。学習。神経回路。人間の問題解決。そして、知能そのもの。
当時のAI研究者たちの周囲には、数年から十数年もあれば、機械は人間の知的能力の多くに追いつくのではないかという楽観があった。チェスも、翻訳も、定理証明も、いったん計算機の問題として置き直せば、あとは速度と工夫の問題に見えたのである。
もちろん、その見積もりは甘かった。
だが、その夏の彼らに、未来の失敗まで見通せというのも酷である。
未知の学問を始める者に、最初から正確な困難の見積もりなどできない。
もし本当の難しさを完全に知っていたなら、誰も漕ぎ出さなかったかもしれない。

三 動く機械——論理理論機

夏の議論の中で、空気を変えた出来事があった。
ニューウェルとサイモンが、ただの構想ではなく、すでに動くプログラムを持ってきたのである。
その名は、論理理論機。Logic Theorist。
RANDのプログラマー、クリフ・ショーとの共同作品である。二人は、それがラッセルとホワイトヘッドの『プリンキピア・マテマティカ』に記された定理を、自力で証明できると説明した。人間の数学者のように、目標を見定め、使えそうな公理や既知の定理を選び、証明への道筋を探る。
もちろん、それは人間の知性そのものではなかった。だが単なる計算でもなかった。すべてを力任せに試すのではなく、見込みのありそうな道を優先する。迷いながら、しかし無闇には迷わない。
そこに、後のAI研究が『探索』や『ヒューリスティック』と呼ぶことになるものの原型があった。
部屋の中にいた者たちは、思い知らされた。
人工知能は、もう言葉だけの夢ではない。
小さく、頼りなく、限定された形ではあっても、すでに動き始めている。
研究の世界には、永遠の真理がある。『議論する者』と『動くものを持参する者』では、後者が常に勝つ。机上で論じられた美しい構想と、実際に動く小さな実装とでは、後者の方がはるかに重い。これを、サイモンとニューウェルは、ダートマスの夏に、皆に教えたのである。
論理理論機は、後に最初期のAIプログラムの代表として語られることになる。
サイモンは後年、心理学・経済学・計算機科学の境界領域で業績を重ね、一九七八年に、ノーベル経済学賞を受賞することになる。AI研究と隣接する分野で、ノーベル賞を受けた、最初の人物である。
ニューウェルは、その生涯をかけて、人間の認知の数学的模型を追求した。彼の遺した『Soar』なる認知体系は、AI研究の重要な遺産の一つとなった。
二人がダートマスの夏へ持ち込んだものは、完成品ではなかった。
それは、灯火だった。
小さく、揺れやすく、風に消えそうな火。
だが、暗い海の上では、その小さな火こそが船の方角を教えることがある。

四 名前が旗になる瞬間

ダートマスの夏における最大の発明は、プログラムではなく、名前だったのかもしれない。
ただし、その名前は、一九五六年の芝生の上で突然生まれたのではない。
Artificial Intelligence。
人工知能。
その言葉は、前年の研究計画書にすでに書かれていた。
命名の瞬間を一つの場面として切り取ることはできない。マッカーシーがどの机で、どの時刻に、その二語を最初に並べたのか。誰とどんな会話を交わしたのか。史料は、映画のようには残っていない。
だが、歴史には、言葉が生まれる瞬間と、その言葉が旗になる瞬間がある。
一九五五年、紙の上に置かれた『人工知能』という語は、一九五六年の夏、研究者たちが実際に集まったことで、共同体の旗になった。
新しい学問には、しばしば挑発的な名前が必要である。
名前は旗である。大義名分である。
まだ領土がなくても、旗が立つと人はそこへ集まりはじめる。集まった人々が境界を描き、方法を作り、論文を書き、学生を育てる。そのうちに、最初は空虚にも見えた名が、本当に中身を持ちはじめる。
人工知能という言葉は、まさにそのような旗だった。
ダートマスの夏に、知能の謎が解けたわけではない。機械が人間のように考え始めたわけでもない。
だが、このとき『AI』と呼ばれる場所ができた。
半世紀以上後、その二文字は世界中の新聞に、企業の看板に、大学の講義名に、政策文書に、そして人々の手の中の小さな画面に現れることになる。
一九五五年に名が書かれ、一九五六年の夏に旗が立った。
地図は、まだなかった。

五 櫓も舵もなき船

時は、一九五六年のダートマスから、遥か昔へ遡る。
一七七一年三月、江戸。小塚原刑場。
罪人の腑分けが行われていた。立ち会っていた杉田玄白、前野良沢、中川淳庵らは、オランダ語の解剖学書『ターヘル・アナトミア』の図と、目の前に開かれた人体とを見比べた。
図が、合っている。
従来の医学書に描かれていた人体図とは違う。目の前の臓器と、オランダの書物の図が、驚くほど一致している。
ならば、この書物は訳さねばならない。
そう思うのは自然だった。
問題は、訳すための道具が、あまりにも足りなかったことである。
一同の中では、前野良沢が最も深くオランダ語を学んでいた。だが、今日のような十分な辞書も、整った翻訳環境もない。玄白たちは、一語ずつ意味を探り、図と本文を照らし合わせ、既知の語から未知の語を推測しなければならなかった。
あるのは、一冊の蘭書と、目の前で見た事実と、これを日本語に移さねばならないという切迫した思いだった。
後に杉田玄白は、その心境を『蘭学事始』に、櫓も舵もない船で大海へ乗り出すようだった、と記した。
櫓もない。舵もない。
それでも大海へ出た。
彼らは一語ずつ推測し、一語ずつ漢語を当て、人体の名を作り、文を組み立てていった。訳せるから始めたのではない。始めたから、少しずつ訳せるようになったのである。
一七七四年、『解体新書』が世に出た。
それは、日本における西洋医学受容の大きな転機となり、蘭学の興隆を象徴する書物の一つとなった。
ダートマスの夏も、どこか似ている。
集まった研究者たちは、明確な結論には至らなかった。各々が異なる方法論を主張し、知能とは何かについての合意すら得られなかった。参加者は夏の数週間を行き来し、それぞれの大学や研究所へ戻っていった。
何かが定まったわけではない。何かが解けたわけでもない。
ただ、彼らは、漕ぎ出したのである。
櫓も舵もなき船で。

六 三つの本山

会議の真の遺産は、その後の数年間に、徐々に明らかになっていった。
マッカーシーは、ダートマスを離れ、やがてマサチューセッツ工科大学へ移った。ミンスキーもMITで研究を進め、のちに同校のAI研究を代表する拠点が形づくられていく。さらにマッカーシーはスタンフォード大学へ移り、スタンフォード人工知能研究所を創設する。
ニューウェルとサイモンは、カーネギー工科大学、後のカーネギー・メロン大学を本拠地とした。彼らの研究は、AIと認知科学の交差点に大きな系譜を残した。
——スタンフォード。MIT。カーネギー・メロン。
この三つは、その後の米国AI研究を代表する主要な拠点となった。
もちろん、AIの歴史がこの三校だけでできているわけではない。エディンバラ、トロント、モントリオール、ベル研究所、IBM、バークレー、そして後の企業研究所。火は、世界の各地で燃えていた。
それでも、ダートマスの夏に集った人々から、いくつもの研究拠点と弟子の系譜が伸びていったことは確かである。
蘭学もまた、人を介して広がった。玄白や良沢らの仕事ののち、蘭学は多くの家塾や私塾へ受け継がれ、やがて緒方洪庵の適塾のような学びの場を生み、その先には福沢諭吉ら近代日本の知の担い手が現れる。
学問は、人を介して、伝わっていく。
本山が立ち、そこから弟子が出て、弟子がまた別の場所で火を守る。
その連鎖の中で、最初の小さな種は、巨大な森へと育っていく。
ダートマスの夏は、その種が蒔かれた瞬間だったのである。

終 二十年後

ダートマスの夏を境に、研究者たちは大胆な見通しを、世に向けて語りはじめた。
その楽観の象徴として、後年よく引かれる言葉がある。一九六五年、サイモンはこう言い切った。
『二十年以内に、機械は人間ができる仕事のすべてをこなすようになる』
マッカーシーらも、似たような楽観を共有していた。
ダートマスの夏から二十年——すなわち一九七六年。
海は、まだ暗かった。
機械翻訳は文脈の深みに沈み、定理証明は小さな世界の外へ出ると苦しみ、ロボットは現実世界の複雑さにつまずいた。チェスも、会話も、視覚も、常識も、人間には当たり前に見えることほど機械には難しかった。AI研究は、深刻な冬の時代に入っていた。資金は枯渇し、学生は他分野へ流れ、夏の夢は、まだほとんど何一つ、実現していなかった。
ニューウェルもサイモンも、ミンスキーもマッカーシーも、皆、まだ生きていた。皆、あの夏に語った楽観が、いかに無謀なものであったかを、痛感していた。
だが、彼らのうち、誰一人として、研究をやめなかった。
櫓も舵もなき船は、まだ大海の真ん中にあった。岸はまだ、見えていない。
それでも、漕ぐのである。漕ぎ続けるのである。彼らは、そう決めていた。
漕ぎ続けた者の見る景色を、私たちは、半世紀後、彼ら自身よりも、よりよく知ることになる。
第二章 了

第三章 蝋の翼

フランク・ローゼンブラットの肖像(一九二八—一九七一)

一 ブロンクスの少年たち

一九四〇年代、ニューヨーク市ブロンクス区。
地下鉄の終着駅近くに、ひとつの高校があった。ブロンクス科学高校。市内から選抜された、数学と科学に異様なほどの適性を持つ少年少女たちが集まる学校である。
その校舎に、二人の少年がいた。
一人は、フランク・ローゼンブラット。鋭い目をした、熱を帯びた少年だった。何かに取り憑かれると、相手が疲れるまで話し続ける。論理だけでなく、身振りや声の抑揚まで使って、自分の見ている未来を人に見せようとするところがあった。
もう一人は、マービン・ミンスキー。早口で、発想が跳ぶ。だがその跳躍は、しばしば周囲が追いつけないほど遠く、正確な場所へ着地した。
二人がどの程度親しく言葉を交わしたのか、詳しい記録は残っていない。同じ廊下を歩いたかもしれない。同じ教師の声を聞いたかもしれない。図書館の同じ棚に手を伸ばしたことがあったかもしれない。
確かなのは、ただひとつ。
後に人工知能の歴史を二つに裂くことになる二人が、若い日に、同じ学校の空気を吸っていたということである。
ローゼンブラットは、脳に似た機械を信じる側へ進む。
ミンスキーは、やがてその限界を告げる側へ立つ。
もちろん、その未来はまだ誰にも見えていない。十代の彼らにあったのは、ただ、科学への飢えと、世界を説明できるはずだという若者特有の傲慢さだけだった。
悲劇は、まだ始まっていなかった。

二 コーネルの蝋翼

一九五七年、ニューヨーク州バッファロー近郊。
コーネル航空研究所の一室で、ローゼンブラットは机に向かっていた。二十八歳。コーネル大学で心理学の博士号を得た彼は、知覚と脳の仕組みに深く魅せられていた。
机の上に広げているのは、紙と鉛筆。図には、いくつもの円と矢印が描かれている。
それは、人間の脳の神経細胞を、数式の形に置き換えたものだった。
一つの細胞が、複数の入力信号を受け取り、それぞれに『重み』をかけて足し合わせ、ある閾値を超えれば、次の細胞に信号を送る——という、最も単純な仕掛け。
これだけの仕掛けでも、『重み』を正しく調整することができれば、機械は何かを『学習』する。光のパターンを認識する。文字を見分ける。音を聞き分ける——そういったことが、原理的には可能なはずだ、と彼は信じていた。
ローゼンブラットは、その仕組みに名前を与えた。
『パーセプトロン(Perceptron)』
知覚するもの、という意味のラテン語から取った造語である。
それは、まだ粗末な翼だった。羽根は少なく、骨組みは単純で、空のすべてを飛ぶにはあまりにも頼りない。だが、それでも翼であることに変わりはなかった。
この頃、人工知能の研究者の多くは、人間の推論を記号と規則で表そうとしていた。論理、探索、定理証明。知性とは、明示された規則を操作することだと考えられていた。
ローゼンブラットの直感は違っていた。
知能は、最初から規則として与えられるものではない。
世界に触れ、誤り、修正されるうちに、内側に形を作っていくものではないか。
彼は、自分が空に飛び立つ翼を、いま、組み立てているのだと感じた。それは、ダイダロスが息子イカロスのために組み上げた、蝋と羽根の翼にも似た、美しい構造をしていた。
蝋の翼は、組み立てさえすれば、本当に空を飛ぶ。
問題は、どこまで飛んでよいか、を、知らないことだった。

三 新聞の予言

一九五八年七月。
パーセプトロンは、研究室の外へ出て、新聞の見出しになった。
ニューヨーク・タイムズは、米海軍が『自ら学び、賢くなる』新しい計算機の胚胎を披露した、と大きく報じた。
その記事は、未来を大胆に先取りした。
将来、この種の機械は、歩き、話し、見て、書き、自己を再生産し、さらには自らの存在を意識するかもしれない——。
重要なのは、この壮大な未来像を、そのままローゼンブラット本人の逐語的な予言として読むべきではない、ということである。そこには、米海軍の期待も、新聞の熱気も、当時の科学技術に向けられた夢も混ざっていた。
それでも、ローゼンブラット自身がパーセプトロンの可能性を大胆に信じていたことは確かだった。
当時できたことは、きわめて限られていた。
単純なパターンを区別する。
例を与えられ、重みを変える。
それだけである。
だが、それまで人間が一つ一つ規則を書いていた機械が、例から自分の内部を変える。
その一点だけでも、十分に新しかった。
新聞は、粗末な翼の先に、まだ存在しない空を見た。
ローゼンブラットもまた、その空を見ていた。
彼は、翼を隠さなかった。
むしろ、それを高く掲げた。
見よ、機械は学ぶのだ、と。
ダイダロスは、息子イカロスに告げたという。
高く飛びすぎてはならない。太陽に近づけば、蝋が溶ける。
低く飛びすぎてもならない。海に近づけば、羽根が濡れる。
しかしイカロスは、飛べることそのものに酔った。
空があることを知った者に、ほどほどの高さを守れと言うのは難しい。
ローゼンブラットもまた、飛んだ。
そして、その飛翔は、時代よりあまりに早かった。

四 マークI

一九五〇年代の終わり、ニューヨーク州バッファロー。コーネル航空研究所。
ローゼンブラットの構想は、紙の上だけのものではなくなった。
部屋には、四百個の光センサーを並べた二十×二十の受光面と、多数の配線、可変抵抗器からなる装置が組み上げられつつあった。
名前は『マークI・パーセプトロン』。
最初期の、本格的なニューラルネットワーク・マシンである。実機が報道陣の前で公開されるのは、一九六〇年のことになる。
それは、今日の目から見れば原始的な装置だった。
だが当時、それは未来そのものに見えた。
機械は、単純な視覚パターンの違いを、例から学ぶことができた。人間が一つ一つの判定規則を書き込むのではない。入力と正解を重ねるうちに、内部の重みが変わっていく。
学習する機械。
その言葉が、実物の重みを持った。
米国海軍が研究を支援した。
新聞記者たちが集まった。
研究室には、若い学生たちが集まった。
ローゼンブラットは、熱心な指導者だったという。学生の話を聞き、時間を惜しまず議論し、自分の熱を分け与えるように研究を語った。
彼は華やかで、時に大げさで、しかし人を惹きつけた。
その一方で、彼には研究室の外にもう一つの場所があった。
海である。
ローゼンブラットはヨットを愛した。
帆を張り、風を読み、水面を滑る。
風が合えば、船は驚くほど軽やかに進む。自分の力ではなく、見えない流れを受けて進む感覚が、彼を解放したのかもしれない。
帆もまた、ある意味、蝋の翼に似ている。
風次第で、どこまでも飛ぶことができる。
だが、嵐が来れば、たちまち海に落ちる。

五 限界の書

一九六九年、マサチューセッツ州ケンブリッジ。
MITのマービン・ミンスキーとシーモア・パパートは、一冊の本を出版した。
題は『Perceptrons』。
それは、パーセプトロンを感情的に断罪する告発書ではなかった。
むしろ、ある種のパーセプトロンが何をでき、何をできないのかを、数学の言葉で徹底的に調べた本だった。
その中で示された限界の代表例として、後世とりわけ有名になったものがある。
排他的論理和。
XOR。
二つの入力のうち、片方だけが真ならば真。両方とも真、あるいは両方とも偽ならば偽。
この配置は、一本の直線では二つに分けられない。
単層の線形パーセプトロンは、世界を一本の直線、あるいは高次元の超平面で分ける。そのため、XORのように線形分離できない問題を表現できない。
この数学的指摘は、正しかった。
だが、歴史は数学だけでは動かない。
多層にすれば表現できる問題が増えることは知られていた。問題は、当時、その多層ネットワークを効率よく学習させる方法、十分な計算資源、十分なデータが揃っていなかったことである。
そのため『Perceptrons』は、のちにニューラルネット研究の停滞を象徴する本として語られるようになった。
しかし、冬を一冊の本だけに帰すことはできない。
過大な期待。
貧弱な計算資源。
限られた実験結果。
研究資金の方向転換。
別のAI手法の台頭。
いくつもの要因が重なっていた。
歴史は、ときに一人の悪役を欲しがる。
だが、現実の冬には、一人の犯人はいない。
それでも、ローゼンブラットの翼がこの時代に飛び続けられなかったことは確かだった。
彼は、まだ四十代の初めであった。

六 七月十一日の海

一九七一年七月十一日。
メリーランド州、チェサピーク湾。
その日は、ローゼンブラットの四十三歳の誕生日だった。
彼はヨットで海に出た。穏やかな一日になるはずだった。だが、風は変わる。水面は荒れる。船は、思いがけない角度で傾く。
ボートは転覆した。
事故の詳しい状況は、伝わっていない。
ローゼンブラットは、戻らなかった。
七月十一日に生まれ、七月十一日に海で命を落とした。
この一致を、単なる偶然として片づけることはできる。事実としては、そうするべきかもしれない。彼の死は、公式には事故とされている。それは、はっきり書いておかなければならない。
だが、歴史には、事実の奥で神話の形を取る瞬間がある。
蝋の翼で空を目指した男が、海に落ちて死ぬ。
イカロスの物語は、ここで静かに影を落とす。ローゼンブラットが太陽に近づきすぎたのか、それとも時代の太陽があまりに早く彼の翼を焼いたのか。それは分からない。
彼は、間違っていたのだろうか。
ある意味では、間違っていた。パーセプトロンは、すぐに歩きも話しもせず、意識も持たなかった。
だが別の意味では、彼は正しかった。
機械は、経験から学ぶ。世界を見て、重みを変える。明示された規則ではなく、無数の例から内側に形を作る。
その直感は、彼の死後も消えなかった。ただ、表舞台から遠ざかり、冷たい海の底へ沈んだだけだった。

終 半世紀ののち

一九八六年、十月。雑誌『ネイチャー』。
デイヴィッド・ラメルハート、ジェフリー・ヒントン、ロナルド・ウィリアムズの三人による、一篇の論文が掲載された。
題は『誤差逆伝播による表現の学習』。
誤差を出力側から入力側へ送り返し、各結合の重みを少しずつ修正する。
バックプロパゲーション。
この考え方そのものには先行研究があった。だが、一九八六年の論文は、多層の神経網が内部表現を学べることを鮮やかに示し、この方法をニューラルネット研究の中心へ押し戻した。
かつて単層パーセプトロンが越えられなかったXORは、多層ネットワークなら越えられる。
より複雑な問題へ向かう道も、見え始めた。
ただし、これで一夜にして冬が終わったわけではない。
計算機はまだ非力だった。
データも足りなかった。
本当の夜明けまでには、さらに時間が必要だった。
そして、二〇一二年——AlexNet。
さらに、二〇二二年——ChatGPT。
一九五八年の新聞が夢想した『見て、学び、書き、話す機械』のうち、いくつかは、別々の技術を組み合わせながら現実になった。
自己複製も、自己意識も、なお慎重に扱うべき問いである。
だが、少なくとも、機械が例から学び、画像を見分け、文章を書き、人間と会話する世界は到来した。
ローゼンブラットの未来像は、すべて正しかったのではない。
ただ、その方向のいくつかは、半世紀ほど早く見えていた。
蝋の翼は、一度溶けた。
海に落ちた。
人々は、それを失敗の物語として語った。
だが後の時代は、同じ夢を別の材料で作り直した。
より強い計算機。
より大きなデータ。
より深い層。
より巧みな学習法。
蝋と羽根でできていた翼は、シリコンと数式と電力で組み直された。
もしローゼンブラットが、半世紀後の世界を見たならば、何と言っただろうか。
おそらく、少し大げさに笑ったかもしれない。
ほら、と。
私が見ていた空は、やはり、あったではないか、と。
七月十一日の海に落ちた翼は、消えたのではなかった。
長いあいだ沈んでいただけである。
そして二十一世紀、かつて蝋で作られたその翼は、世界中の機械の内側で、別の素材となって再び羽ばたき始めた。
第三章 了

第四章 隠れた灯火

冬と蘭学者たち(一九六九—一九八六)

一 冬の到来

一九七三年、英国。
一通の報告書が、AI研究界に冷たい風を送り込んだ。
報告書の作成者は、ジェームズ・ライトヒル卿。流体力学の権威であり、AIの専門家として研究を率いてきた人物ではなかった。英国の科学研究会議は、彼に人工知能研究の現状を評価させた。
ライトヒルの評価は厳しかった。
AI研究は、初期に掲げた大きな約束に比べ、現実の成果が十分ではない——という趣旨である。
この報告書ののち、英国ではAI研究への公的支援が大きく縮小された。
米国でも、ARPAの資金は次第に、より明確な実用成果を求めるようになっていく。
ダートマスの夏から十七年。
初期の研究者たちが語った楽観の多くは、期限を過ぎても実現しなかった。
機械翻訳は、文脈の壁にぶつかった。
画像認識は、現実世界の複雑さに苦しんだ。
音声認識も、限られた条件の外へ出ると難しかった。
神経網研究も、単純なパーセプトロンの限界、計算資源の不足、研究潮流の変化の中で、主流から遠ざかっていた。
AIの最初の冬が、到来したのである。
ただし、冬は一日で来るものではない。
一冊の報告書、一冊の本、一度の失敗だけで世界が凍るわけではない。
過大な期待があり、期待外れがあり、資金の方向転換があり、研究者の関心の移動がある。
それらが重なったとき、季節は変わる。
ダートマスの夏は、遠くなっていた。

二 朱子学の春

冬の中にも、しかし、別の場所では、思いがけぬ春が訪れていた。
カリフォルニア州、スタンフォード大学。
ノーベル賞学者ジョシュア・レーダーバーグと、計算機学者エドワード・ファイゲンバウムらは、ある大胆な試みに取り組んでいた。
人間の専門知を、機械の中へ移す。
万能の知能をいきなり作るのではない。
狭い領域を選び、その領域の専門家が使う知識や判断の筋道を、機械が利用できる形にする。
最初の代表的な試みは、有機化学の領域で行われた。
プログラムの名前は『DENDRAL』。
質量分析などのデータを手がかりに、候補となる有機化合物の構造を推定する。
そこでは、化学者の専門知識と、探索の仕組みが結びつけられていた。
それは万能の知能ではなかった。
会話もできず、世界の常識も知らず、詩も書けない。
しかし、狭い領域では強かった。
DENDRALは、AIに一つの教訓を与えた。
広すぎる知能を夢見るより、狭い領域で深い知識を持たせた方が、機械は役に立つのではないか。
ファイゲンバウムは、こうした発想を『知識工学』という言葉で押し広げていく。
本書の比喩でいえば、これは徳川時代の朱子学に少し似ている。
世界には秩序があり、その秩序は言葉にできる。
学ぶべき知識を整理し、伝え、適切な場面で適用する。
藩校や私塾で、若者たちが古典の言葉を通じて世界の秩序を学ぼうとしたように、記号主義AIは、人間の知識を明示的な形で機械へ渡そうとした。
もちろん、朱子学とAIが同じ思想だったわけではない。
ここで似ているのは、知を記述し、体系化し、継承可能な形にしようとする身振りである。
その信念が、一九七〇年代から八〇年代前半にかけて、人工知能研究の大きな潮流となっていく。

三 医療現場のMYCIN

スタンフォード医学部。一人の若き医師が、ファイゲンバウムの方式に注目した。エドワード・ショートリフ。当時、二十代後半。
ショートリフは、ある専門領域に、この方式を応用しようと考えた。
血液感染症の診断と、抗生物質の選択。
これは、当時の病院で、しばしば判断が分かれる難しい領域であった。医師によって診断が違い、処方される薬も違った。患者の命を左右する、経験がものを言う繊細な判断の連続であった。
ショートリフは、各分野の専門医を訪ね、彼らの判断の流れを、丁寧に聞き出した。
『もし、患者の血液からこの種の細菌が検出され、かつ、患者の年齢がこの範囲なら、抗生物質Xを処方する』
『もし、患者にこの種の既往歴があり、かつ、血液検査でこの値が高ければ、抗生物質Yを優先する』
——こうしたルールを、約五百個、彼は集めた。
それらを、機械の中に組み込んだ。
機械の名前は『MYCIN』。MYCINは、単純な『はい』か『いいえ』だけで動く機械ではなかった。不確実性を扱うために、確信度という考え方も取り入れた。医師の判断がいつも百パーセントの断定ではないように、機械もまた、可能性の濃淡を扱おうとしたのである。
一九七〇年代後半、初期の評価実験が行われた。MYCINの判断と、専門医の判断とを、比較する。評価実験では、MYCINは専門医に匹敵する結果を示した。
これは衝撃だった。
機械が、医師の判断に近づいたのである。
もっとも、MYCINがそのまま病院で広く使われたわけではない。責任の所在、入力の手間、医療制度との接続、現場の信頼。現実の医療には、診断精度だけでは越えられない壁があった。
それでも、MYCINの意味は大きかった。
機械は、専門家の知識を規則として持てば、現実の問題に答えられるかもしれない。その期待は、産業界へ広がっていく。
同じ頃、DEC(デジタル・イクイップメント社)では、『XCON』と呼ばれるエキスパートシステムが、計算機の構成設定の自動化に投入され、年間数千万ドルの経費削減をもたらしていた。
人工知能は、ふたたび実用の言葉で語られはじめた。
ファイゲンバウムは言った。
知識こそが力である。
それは、エキスパートシステム時代の標語となった。
朱子学の春は、いまや満開に見えた。

四 第五世代の夢

その春は、太平洋を渡って日本にも届いた。
一九八二年、東京。
一九八二年四月、通商産業省(通産省)の主導により、ある国家プロジェクトが始動した。
『第五世代コンピュータ計画』。
この名は、計算機の歴史を、五世代に分けた区分から来ている。真空管の第一世代、トランジスタの第二世代、集積回路の第三世代、超大規模集積回路の第四世代——そして、人工知能を備えた、第五世代。
日本が、世界の計算機産業を、AIによって一気に追い越そうという、壮大な国家構想であった。
予算は、十年間で総額およそ五百四十億円。当時としては、米国・欧州が震えるほどの規模であった。
本部となる組織は、『新世代コンピュータ技術開発機構』——通称ICOT。東京・三田の地に、研究拠点が置かれた。
そのプロジェクトの中心に立ったのが、渕一博、四十六歳。電子技術総合研究所より派遣された、論理プログラミングの専門家であった。
渕の仕事は、研究者の仕事であると同時に、翻訳者の仕事でもあった。研究者の言葉を行政へ伝え、行政の構想を企業へ伝え、企業の技術者たちを一つの大きな目標へ向かわせる。日本中の若い計算機科学者、企業研究者、技術者たちが、ICOTへ集まった。
当時、そこには本物の熱があった。
日本が、次の計算機の時代を主導する。
単なる数値計算の機械ではなく、知識を扱う機械を作る。
論理によって推論し、自然言語を扱い、専門家のように答えるコンピュータを実現する。
海外は警戒した。米国は警戒した。レーガン政権下で『戦略的計算機構想』が立ち上げられ、欧州では『エスプリ』『アルベイ』といった対抗プロジェクトが次々と起こされた。日本の挑戦は、それだけ重く受け止められたのである。ICOTが選んだ中核技術は、論理プログラミングだった。
一九七二年に欧州で開発されたPrologに代表される論理型言語を発展させ、並列推論機械の上で大規模に動かす。人間が手続きの細部を書くのではなく、事実と規則を与え、機械が推論によって答えを導く。
これもまた、朱子学的な世界観に立っていた。
世界は、論理によって記述でき、論理によって推論できる、と。
渕一博は、この理想を、誰よりも信じていた。

五 規則の森

だが、春は続かない。
一九八〇年代後半、エキスパートシステムの限界が見えはじめた。
当初は、規則を増やせば増やすほど賢くなるように思われた。だが実際には、増えるほど絡み合った。専門家Aの規則と、専門家Bの規則が衝突する。新しい知識を入れると、古い規則の一部が壊れる。例外を加えると、さらに例外が必要になる。
知識は、書物の棚のように整然とは積み上がらなかった。
なにより問題だったのは、常識だった。
専門領域の知識ならば、まだ聞き出せる。医師に、化学者に、技術者に尋ねればよい。しかし、人間が日常で使っている常識は、あまりに広く、あまりに暗黙の裡にあった。
『水は、こぼれれば、下に落ちる』『人は、食事をしないと、お腹が空く』『夜になれば、太陽は沈み、月が昇る』——こうしたことの一つ一つを、ルールの形に書き出していこうとした研究もあった。だが、世界の常識を全部書き出すのは、無限に近い作業であった。人間にとって当たり前すぎることほど、規則に書き出すのは難しかった。
エキスパートシステムは、狭い部屋では賢かった。だが、廊下へ出ると迷った。建物の外へ出ると、ほとんど何も知らなかった。
やがて、専用のLISPマシン市場も崩れ始める。高価な専用機は、急速に性能を上げる汎用ワークステーションに追い上げられた。企業は、人工知能という言葉に再び慎重になっていった。
第五世代計画もまた、当初の壮大な目標をそのまま実現することはできなかった。
この物語の時間を少し先へ進めれば、一九九二年、十年の計画期間を終え、ICOTは一区切りを迎える。並列推論機械、論理プログラミング処理系、知識処理の研究など、成果は確かにあった。しかし、『知能を持つ第五世代コンピュータ』が社会を一変させたわけではなかった。
渕一博は、評価会の場で、淡々とプロジェクトの結果を報告した。
彼は、自分の責任を、回避しなかった。
当初の壮大な構想に照らせば、達成はその一部にとどまった。その評価を、彼は正面から受け止めた。
だが、彼が最後まで指し示したものがある。
ここで育った技術者たちが、これからの日本の情報技術を支えていく、という確信である。
後にこの確信は現実となる。ICOT出身の研究者たちは、日本の計算機科学の中核に散らばり、それぞれの場所で、長い貢献を続けることになるのである。
朱子学は、それ自体としては力を失っても、その教養を身につけた人々が、明治の改革を支えた。
しかし、AIの進歩にとっては、第二の冬が、訪れていたのである。

六 蘭学者の灯火

冬の中で、別の場所では、ひそかに、別の灯火が燃え続けていた。
スコットランド、エディンバラ。
ある研究室に、若き心理学者がいた。ジェフリー・ヒントン。一九四七年、ロンドン生まれ。エディンバラ大学で、人間の認知を計算機の中に再現する研究をしていた。
彼は、神経網——ローゼンブラットが残した、あの『時代遅れ』とされた研究分野——を、頑なに信じ続けていた。
ロンドンに生まれ、エディンバラで人工知能を学んだ彼は、神経網と分散表現に強く惹かれていた。知識は、必ずしも一つ一つの記号として書かれているのではない。多数の重みの中に、分散して宿るのではないか。
これは、記号主義の正統から見れば異端だった。
意味は辞書の項目のように明示されるものではない。
規則は人間が全部書くものではない。
機械は、例を通じて、自ら内部の形を変える。
ローゼンブラットの火は、まだ消えていなかった。
ヒントンは、何度か職を変えた。エディンバラから、カリフォルニア大学サンディエゴ校へ。そこで、デイヴィッド・ラメルハートら『PDPグループ』と呼ばれる若き研究者たちと出会った。次にカーネギー・メロン大学へ。最後にカナダ、トロント大学へ。
トロントに移った理由について、後年、彼は半分冗談まじりにこう語っている。
『米国にいると、AI研究の資金は、軍事用途を意識せねばならない。私はそれが嫌いだったから、軍事色の薄いカナダに来た』
同じころ、デイヴィッド・ラメルハートらは、カリフォルニアで並列分散処理、PDPと呼ばれる考え方を育てていた。人間の認知を、中央の司令塔ではなく、多数の単純な単位の相互作用として捉える見方である。
フランスでは、ヤン・ルカンが神経網に取り組んでいた。彼は後に、畳み込み神経網を手書き数字認識へ応用し、画像認識の流れに大きな足跡を残すことになる。
さらに若い世代として、カナダのヨシュア・ベンジオが続く。
彼ら三人は、後に『コネクショニズム(接続主義)の三巨頭』と呼ばれることになる。
しかし、一九八〇年代当時、AIの第二の冬の中で、彼らは、研究界の主流からは、明らかに外れた場所にいた。
三人は、互いに論文を読み、ときに同じ学会で顔を合わせた。彼らの間に、ゆるやかな同志意識があった。
『神経網は、まだ終わっていない』
それが、三人の、ささやかな信念であった。
歴史を動かす灯火は、しばしばその程度の小ささで始まる。

終 一九八六年、冬の終わり

一九八六年十月。
雑誌『ネイチャー』に、一篇の論文が掲載された。
デイヴィッド・ラメルハート、ジェフリー・ヒントン、ロナルド・ウィリアムズによる『誤差逆伝播による表現の学習』。
バックプロパゲーションという考え方には、それ以前から先行する研究があった。
だが、この論文は、多層の神経網を学習させる方法として、その力を鮮やかに示した。
出力が間違う。
その誤差を、後ろの層から前の層へ送り返す。
どの重みがどれだけ間違いに寄与したかを計算し、少しずつ修正する。
それを何度も繰り返す。
この仕組みによって、多層の神経網を実際に学習させる道が、大きく開かれた。
かつて単層パーセプトロンが越えられなかったXORは、もはや小さな例題になった。
より複雑な問題へ向かうための扉が、わずかに開いたのである。
もちろん、これで冬が終わったわけではない。
計算機はまだ非力だった。
データも足りなかった。
神経網は再び注目されたものの、世界を変えるには、まだ多くの時間が必要だった。
本当の夜明けまでには、さらに二十六年を待たねばならない。
二〇一二年。
AlexNet。
だが、一九八六年の時点で、雪の下にはすでに春の息吹があった。
朱子学の大伽藍が揺らぎ、正統の言葉が力を失いはじめたころ、蘭学者たちは小さな灯を守っていた。
彼らはまだ勝者ではなかった。
世界はまだ、その灯の意味を知らなかった。
それでも、火は消えていなかった。
人工知能の冬の奥で、ひそかに燃え続けていたその灯火が、やがて二十一世紀の世界を照らすことになる。
第四章 了

第五章 黒船とサムライ

統計革命とディープブルー(一九九〇—一九九七)

一 開国前夜

一九九〇年代の幕開け、人工知能研究は再び深い霧の中にあった。
エキスパートシステムの春は過ぎ去っていた。知識を規則として書き込めば機械は賢くなる——その信念は、常識という底なし沼の前で力を失っていた。
日本の第五世代コンピュータ計画も、壮大な理想を掲げながら、約束された未来をそのまま実現することはできなかった。論理プログラミングと推論機械の夢は、確かな技術と人材を残しながらも、時代全体を塗り替えるには至らなかった。
神経網は、一九八六年の誤差逆伝播研究によって息を吹き返したかに見えた。だが、まだ主流ではなかった。計算機は非力で、データは少なく、世間はニューラルネットワークという言葉に半信半疑だった。
人工知能という看板の外側で、別の研究領域が力を増していた。
機械学習。
パターン認識。
統計的推定。
データ解析。
それらは単なる『AIの偽名』ではない。それぞれに独自の歴史と方法を持つ学問である。
だが結果として、AIが大きな夢を語りにくかった時代にも、その周辺で、次の時代を作る技術は育ち続けた。
この空気は、幕末の日本に少し似ている。
徳川の世は、長く続いた。秩序があり、規則があり、身分があり、言葉があった。だが十九世紀半ば、その秩序の外側から、別の力が近づいていた。誰もが何かが変わると感じていた。だが、何に変わるのかは、まだ見えていなかった。
AI研究も同じだった。
記号と論理の幕府は、まだ倒れてはいない。
しかし、沖合にはすでに黒い影が見えていた。
その影は、蒸気船ではなかった。
数式とデータでできた、知の黒船だった。

二 黒船は数式で来た

嘉永六年六月三日。西暦では一八五三年七月、浦賀沖。
四隻の異国船が現れた。黒く塗られた船体。煙突から上がる黒煙。風がなくとも進む船。帆ではなく、蒸気で海を渡る船である。
ペリー提督は、米国大統領フィルモアの親書を携えていた。
日本は開国し、米国と交わるべし。
幕府は揺れた。
開国か、攘夷か。守るか、変わるか。
黒船は、日本をその場で征服したわけではない。だが、それまでの世界観を打ち砕いた。もはや日本は、自分の内側だけで世界を閉じることはできない。その事実を、黒煙を上げる船が誰の目にも示したのである。
一九九〇年代、AI研究にも、似たような『黒船』が現れていた。
それは、一つの国から来た一隻の船ではない。
互いに異なる海から現れた、複数の船団であった。
ヴラジーミル・ヴァプニクは、統計的学習の境界を研究した。
ジューディア・パールは、不確実性を確率の網で扱い、のちには因果を計算の対象へ押し広げた。
フレデリック・イェリネックは、音声と言語を大量のデータと確率モデルで扱った。
三人は、同じ理念を掲げた同志ではない。
研究対象も、方法も、思想も違う。
だが、彼らの仕事を同じ時代の沖合から眺めると、一つの変化が見えてくる。
世界のすべてを、人間の手で『もしAならばB』という規則に書き直さなくてもよい。
不確実さは、確率として扱える。
分類の境界は、データから学べる。
言葉の並びにも、統計的な規則性がある。
これは、記号主義だけではない別の知能の作り方が、力を持ちはじめたことを意味していた。
蒸気船が、帆船を一夜にして消したわけではない。
統計もまた、論理を一夜にして消したわけではない。
だが、海の進み方は変わり始めていた。

三 境界線と確率の網

ヴラジーミル・ヴァプニクは、ソビエト連邦で長く統計的学習理論を研究していた数学者である。
彼の仕事は、鉄のカーテンの向こう側で育った。やがて米国へ移り、ベル研究所に加わると、その理論は西側の研究社会でも大きな注目を集めていく。
サポートベクターマシン、SVM。
それは、データを分類するための強力な方法だった。
ただ境界線を引くのではない。
二つの集団を、できるだけ大きな余裕をもって分ける境界を探す。
訓練データにぴったり合わせすぎず、未知のデータにも耐える分け方を求める。
機械は、規則を暗記するのではない。
データから、よい境界を学ぶ。
一方、ジューディア・パールは、不確実性の世界を見つめていた。
現実は、真か偽かだけでできていない。
患者の症状は、ある病気の可能性を高めるかもしれないが、別の病気の可能性も残る。
観測された出来事から、見えない原因を推測しなければならない。
パールが体系化したベイジアン・ネットワークは、変数同士の依存関係を確率の網として表現し、不確実な状況で推論するための強力な道具となった。
そして彼の仕事は、さらに因果関係そのものを、明示的に考える方向へ進んでいく。
これは、人工知能にとって大きな転換だった。
それまでの多くのAIは、論理の世界を好んだ。
正しいか、正しくないか。
証明できるか、できないか。
だが現実の知能は、もっと濁っている。
人間は日々、不完全な情報で判断している。
ならば機械もまた、不確実性を扱えなければならない。
SVMとベイジアン・ネットワークは、同じ理論ではない。
それでも、それぞれ別の方向から、AIの重心を広げた。
論理だけでなく、学習へ。
確実性だけでなく、確率へ。
文明開化とは、単に新しい道具が入ることではない。
ものの考え方そのものが増えることである。

四 言語学者を解雇するたびに

統計革命の象徴的な戦場の一つが、音声認識だった。IBMの研究者フレデリック・イェリネックは、チェコ生まれのユダヤ系研究者である。戦争と亡命の影を背負いながら米国へ渡り、やがて音声と言語を統計で扱う研究の中心人物となった。
音声認識は、AIにとって難題だった。
人間の話す音は、曖昧で、崩れ、途切れ、重なる。同じ単語でも人によって発音が違う。雑音も入る。話し言葉は文法どおりに進まない。
従来の発想では、言語学者が文法規則を整え、発音規則を書き、機械にそれを教えるべきだと考えられていた。言葉を理解するには、言葉の構造を明示的に知る必要がある、と。
イェリネックは、その発想を疑った。
機械は、本当に文法を理解しなければならないのか。
大量の音声と文章の対応を見せれば、どの音がどの単語に対応しやすいか、どの単語の次にどの単語が来やすいかを、統計的に学べるのではないか。
隠れマルコフモデル。
Nグラム。
隠れマルコフモデルもNグラムも、言葉の意味を人間のように理解しているわけではない。それまでに現れた語や状態を手がかりに、『次に何が来るのがもっとも自然か』を確率で判断する。つまり、意味そのものではなく、もっともありそうな列を選ぶ仕組みである。言葉を、規則の体系としてではなく、確率の流れとして扱う。こうした確率モデルを使えば、現象を『ある結果が起きる確率はいくつか』という形で記述・予測することができる。
彼は、冗談とも本気ともつかぬ調子で、こう言ったと伝えられる。
『言語学者を一人解雇するたびに、音声認識の性能が上がる』
——言語の規則を諦めることが、言語認識の進歩につながる。
これは、論理に立脚した記号主義AIへの、鋭い挑戦状であった。
明治政府が廃刀令や廃藩置県によって、武士の制度を解体したように、統計革命は、手作業の規則と専門家の直観をAIの中心から降ろしていった。
もちろん、武士が消えたからといって、剣の技や倫理がただちに無意味になったわけではない。同じように、言語学も論理も消えたわけではない。
ただ、新しい時代の中心には、別のものが座った。
データである。

五 チェス界最後のサムライ

時は、一九九〇年代後半。
世界チェス選手権者、ガリー・カスパロフ。
一九六三年、ソ連領アゼルバイジャンのバクー生まれ。十二歳でソ連青年チャンピオン、十七歳で世界青年チャンピオン、二十二歳で当時最年少の世界選手権者——史上最強のチェス棋士の一人と、誰もが認める存在であった。
彼は、二〇〇〇年代に至るまで、世界ランキング一位の座を保ち続けた。十年以上、誰にも破られない、絶対的な王であった。
カスパロフは、ある意味、西郷隆盛に似ていた。
西郷は、明治維新の立役者でありながら、最後には新政府軍と戦い、敗れて命を絶った。彼は、伝統的な武士の価値観——名誉、忠義、潔さ——を、最後まで体現した男であった。新時代の到来を誰よりも早く見抜き、その来航を準備しながら、自らはその新時代に馴染めず、旧き世界とともに散った。
カスパロフは、それとは少し違ったが、構造は似ていた。
彼は、人間のチェスの、最高峰であった。同時に、彼は、知の歴史の中で、ある最後の世代に属していた——『人間が、機械より深く読む』ことが可能であった、最後の世代に。
そして、彼に対して、戦いを挑んでくる機械が、現れた。
カスパロフもまた、単なる旧時代の敗者ではなかった。彼は機械チェスの進歩を軽視していたわけではない。コンピュータを研究にも使い、準備にも取り入れていた。新時代を誰よりも近くで見ていた。
しかし、最後の一線では信じていた。
人間の直観は、機械の計算を超える。世界王者の深い読みは、力任せの探索には敗れない。
そして、彼のその信念に対して、戦いを挑んでくる機械が、現れた。IBMの『ディープブルー』。
チェス専用の計算機械。一秒間に二億手を読み解く、計算の獣であった。
カスパロフは、最初、これを軽く見ていた。
『機械が、人間の直観に勝てるはずがない』と、彼は信じていた。

六 ニューヨーク、十九手

一九九六年二月、フィラデルフィア。
カスパロフとディープブルーの最初の公式マッチが行われた。
第一局、ディープブルーが勝った。
世界王者が、通常の持ち時間の対局でコンピュータに一局を落とした。これは歴史的事件だった。
だが、カスパロフは崩れなかった。第二局以降、彼は機械の弱点を探り、圧力をかけ、揺さぶった。最終結果は、カスパロフの三勝一敗二分け。得点にして四対二。
人間は、まだ勝った。
世界は安堵した。
機械は強い。だが、最強の人間にはまだ届かない。
一年後。
一九九七年五月、ニューヨーク。
再戦の場は、エクイタブル・センター。ディープブルーは改良されていた。より速く、より深く、より巧みに調整されていた。
第一局は、カスパロフが勝った。
しかし第二局で、空気が変わる。
ディープブルーは、カスパロフの予想を超える手を指した。単なる物取りでも、短期的な得でもない。長期的な構想を含むように見える手だった。カスパロフには、それが人間の手に見えた。
彼は敗れた。
そして疑念を抱く。
本当に、この機械は独力で指しているのか。
人間の名手が、どこかで介入しているのではないか。IBM側は否定した。人間の不正な介入を示す証拠はない。だが、重要なのは、証拠の有無だけではなかった。カスパロフの心に、疑念が生じたことそのものが大きかった。
盤上で最も危険なのは、相手の手ではない。
自分の心が揺れることである。
第三局、第四局、第五局は引き分けた。
そして第六局。
カスパロフは、カロ・カン防御を選んだ。堅実で、慎重な序盤である。だがその日の彼には、いつもの攻撃的な覇気がなかった。安全に入り、かえって危険を招いた。
序盤で、彼は致命的な隙を残す。
ディープブルーは、それを見逃さなかった。
十九手。
わずか十九手で、カスパロフは投了した。
世界チェス王者が、複数局の公式マッチで、機械に敗れた。
その瞬間、チェスの歴史だけでなく、人工知能の歴史も変わった。
会見場のカスパロフは、疲れ切っていた。怒りも、疑念も、屈辱もあっただろう。だが、そこにあったのは一人の棋士の敗北だけではなかった。
人間知性の象徴が、機械の前で膝をついた。
そのように世界は受け取ったのである。

終 黒船のあとで

ディープブルーの勝利は、人工知能が人間のように考えたことを意味しない。
それは、重要な点である。
ディープブルーは、人間のようにチェスを愛していたわけではない。美しい手に震えたわけでも、敗北を恐れたわけでもない。ただ膨大な局面を読み、評価し、最善と思われる手を選んだ。
それで、勝った。
この『それで』が、歴史を動かした。
ペリーの黒船も、日本をその場で占領したわけではない。だが、世界の尺度が変わったことを示した。もはや内側の秩序だけでは足りない。外から来た力を認め、取り入れ、作り替えなければならない。
ディープブルーもまた、そうだった。
人間の直観だけでは、もはや十分ではない。
機械の計算力を無視して、知能を語ることはできない。AI研究は、実際に動き、測定され、競争に勝つものへと変わっていく。
同じころ、統計革命も研究室の中で進んでいた。音声認識、文字認識、機械翻訳、情報検索。あらゆる分野で、手作業の規則よりも、大量のデータから学ぶ方法が力を持ちはじめる。
黒船の来航は、明治維新へつながった。
そしてAIにおいても、黒船の衝撃は、新しい体制を生んだ。
論理だけの幕府は終わり、データと確率と計算力の時代が始まった。
カスパロフは、その後も長くディープブルー戦について批判を続けた。彼には、納得できないものが残っていた。敗者の怒りであり、王者の誇りでもあった。
だが、彼はそこで終わらなかった。
やがて彼は、人間と機械が組んで戦う新しいチェスを提唱する。アドバンスト・チェス、あるいはフリースタイル・チェス。人間対機械ではない。人間と機械が協力し、別の人間と機械の組と戦う。
そこで明らかになったのは、興味深い事実だった。
少なくとも初期の大会では、最強の人間が単独で勝つわけでも、最強の機械がそのまま勝つわけでもなかった。勝つのは、機械を最もうまく使う人間だった。人間の判断と、機械の計算。直観と探索。経験とデータ。その組み合わせが、もっとも強かった。
カスパロフは、敗北の中から別の答えを見つけたのである。
機械は、人間の終わりではない。
人間の新しい道具である。
それは、明治日本が欧米の技術を取り入れながら、単なる模倣に終わらず、独自の近代を作ろうとした姿に似ている。
開国の最初の感情は、屈辱である。
だが、屈辱だけでは歴史は終わらない。
その先に、学び、翻訳し、混ぜ合わせ、作り替える時代が来る。
統計革命とディープブルーは、AIにその時代をもたらした。
そして二十一世紀に入ると、データはさらに増え、計算機はさらに速くなり、学習の方法はさらに深くなる。
二〇一二年。AlexNet。
その夜明けは、すでに遠くから近づいていた。
第五章 了

第六章 夜明けの工房

ディープラーニング革命とAlexNet(二〇〇六—二〇一二)

一 春のさきがけ

二〇〇六年、カナダ、トロント大学。
ジェフリー・ヒントンは、五十八歳になっていた。
彼は長いあいだ、神経網研究を守り続けてきた。エディンバラ、カリフォルニア、カーネギー・メロン、そしてトロント。土地を移り、大学を移り、時代の流行が変わっても、彼の信念は変わらなかった。
知能は、必ずしも明示された規則として書かれているのではない。
知識は、記号の棚に整然と並ぶものではない。
それは多数の結びつきの中に、分散して宿る。
だが、世間は長くその考えを信じなかった。
神経網は、古い夢と見なされていた。ローゼンブラットのパーセプトロンは、半世紀近く前に太陽へ近づきすぎ、海へ落ちた翼であった。ヒントンは、その沈んだ翼の破片を拾い集めるように、研究を続けていた。
劇場で未来を語ったローゼンブラットとは違い、ヒントンは工房の職人に近かった。
火が消えないよう、毎朝、炉に薪をくべる者であった。
その彼が、二〇〇六年、一つの重要な道を示す。
深い層を持つ神経網は、当時、学習が難しかった。
勾配が弱くなる問題。
最適化の難しさ。
計算資源とデータの不足。
深くすればよいと分かっていても、うまく学ばせる方法が足りなかった。
ヒントンらは、層を一度に全部訓練するのではなく、一層ずつ事前学習させて積み上げる方法を示した。制限ボルツマン機械を用いた、深い信念ネットワークである。
論文の題は、『深層信念網の高速学習法』。
この仕事は、深いニューラルネットワークへの関心を再び強く呼び戻した。
『ディープラーニング』という言葉そのものを、ヒントンが二〇〇六年に発明したわけではない。
だが、この時期を境に、深い層を学習させる研究の潮流は勢いを増し、やがて『ディープラーニング』という名が時代の旗になる。
それは、十四世紀イタリアのペトラルカに少し似ていた。
ペトラルカは、書庫に眠っていた古代ローマの文献を探し出し、人々に古典の価値を思い出させた。
ヒントンもまた、同じことをしていたのかもしれない。
神経網は死んでいない。
ただ、まだ時代が追いついていないだけだ。
長き中世は、終わろうとしていた。

二 新しい絵具

ルネサンスは、一人の天才だけでは生まれない。
天才には、道具がいる。
工房がいる。
顔料がいる。
注文主がいる。
そして、偶然がいる。
ディープラーニングにも、思いがけない道具が現れた。GPUである。GPUは、もともと人工知能のために生まれたものではない。ゲーム画面を滑らかに描くための部品だった。三次元空間を動く光、影、物体、爆発、煙。それらを一瞬で計算し、画面に映すために、GPUは大量の単純な計算を並列にこなすよう作られていた。
だが——ある時、誰かが気づいた。
『神経網の計算は、本質的に、大量の並列計算ではないか』
各神経細胞が、独立に信号を計算する。それを一斉に並列で処理できれば、訓練の速度は、桁違いに上がる。
神経網の学習は、巨大な行列計算の連続である。無数の重みを掛け、足し、誤差を戻し、また重みを変える。ひとつひとつの計算は単純でも、数は膨大である。
ならば、それをGPUにやらせればよい。
二〇〇七年前後、NVIDIAはCUDAを公開し、GPUを画像描画以外の汎用計算にも使えるようにした。研究者たちは、これを神経網の訓練へ転用し始める。
結果——計算速度は、百倍以上にも、跳ね上がった。
それまで何週間もかかった訓練が、現実的な時間で終わる。
試せなかった大きさのネットワークが、試せる。
諦めていた深さが、届く距離になる。
ゲーマーのために磨かれた部品が、人工知能の工房に持ち込まれたのである。
これは、ルネサンス期の油絵具に似ていた。
それまで、絵画は壁画とテンペラ画が中心であった。だが、十五世紀、ファン・エイク兄弟が油絵具を改良したことで、画家たちは、極めて細密で、何度でも修正可能な絵画を、描けるようになった。レオナルド・ダ・ヴィンチの『モナ・リザ』も、ティツィアーノの豊麗な色彩も、この新技法なくしては、生まれなかった。GPUは、ディープラーニング時代の、油絵具であった。
工房の棚に、新しい絵具が置かれたのである。

三 フェイフェイ・リーの大伽藍

だが、道具だけでは足りなかった。
機械に世界を見せるには、世界そのものが必要だった。
二〇〇〇年代後半、プリンストン大学(のちスタンフォードへ移る)のフェイフェイ・リーは、画像認識の停滞を別の角度から見ていた。問題は、アルゴリズムだけではない。そもそも機械は、世界を十分に見ていないのではないか。
『機械に世界を見せるためには、世界の写真が、足りていない』
これまでの画像認識研究では、せいぜい数万枚程度の画像で機械を訓練していた。だが、人間の小児は、生まれてから数年で、数億枚もの画像を見ている。脳が学習するには、それくらいの量が必要なのではないか——と、彼女は仮説を立てた。
人間の小児は、膨大な量の視覚経験を通じて育つ。犬、猫、椅子、皿、車、木、空、顔、影、雨、雪。世界は、数万枚の訓練画像で足りるような貧しいものではない。
ならば、機械にも世界を見せなければならない。
フェイフェイ・リーは、巨大な画像データベースを構想した。英語の語彙体系WordNetを土台に、物体カテゴリを整理し、インターネット上から画像を集め、人間の手でラベルを付ける。
猫。
犬。
消防車。
灯台。
苺。
椅子。
カワセミ。
飛行機。
トースター。
バイオリン。
その数は、やがて一千万枚を超えた。
この作業を支えたのは、Amazon Mechanical Turkを通じて参加した、世界中の名もなき労働者たちだった。彼らは一枚一枚の画像を見て、それが何であるかを答えた。報酬はわずかであり、名前が歴史に残ることもほとんどない。
だが、その手作業がなければ、のちの革命は起こらなかった。ImageNet。
二〇〇九年に発表されたそのデータベースは、ディープラーニング時代の大伽藍だった。
中世の大聖堂は、一人の建築家だけで建ったのではない。石工、木工、ガラス職人、運搬人、寄進者、祈る人々。世代を超えた無数の手が、尖塔を空へ伸ばした。ImageNetも同じだった。
ひとりの研究者の構想と、世界中の無名の作業者たちの手が、AI研究の巨大な聖堂を築いたのである。
そしてフェイフェイ・リーは、もう一つの仕掛けを用意した。ImageNet Challenge。
世界中の研究室が、自分たちの画像認識システムを持ち寄り、同じデータで精度を競う。誰が世界を最もよく見分けられるのか。その舞台が、用意された。
この競技が、後に、革命の舞台となるのである。

四 ヒントン工房

再び、トロント大学。
ヒントンの研究室には、二人の若い研究者がいた。
アレックス・クリジェフスキー。ウクライナに生まれ、カナダで育った計算機科学者である。寡黙で、実装に強く、計算機の細部を執拗に詰めることができた。彼は理論を飾るより、動くものを作る人間だった。
イリヤ・サツケヴァー。ソ連に生まれ、幼くしてイスラエルへ移り、その後カナダへ渡った。早熟で、抽象的な直観に優れ、深層学習の可能性を深く信じていた。
師であるヒントンは、彼らに静かな信頼を置いていた。
この研究室は、十五世紀フィレンツェのヴェロッキオ工房に似ている。
ヴェロッキオは、彫刻家であり、画家であり、工房の親方だった。そこには若い才能が集まった。なかでも、レオナルド・ダ・ヴィンチがいた。師は弟子に技を教え、弟子は師の技を越えようとした。
ヴェロッキオの『キリストの洗礼』には、若きレオナルドが描いたと伝えられる天使がいる。その天使があまりに美しかったため、ヴェロッキオは以後、絵筆を取らなくなったという逸話が残る。
真偽はともかく、この逸話は工房というものの本質を語っている。
工房とは、師がすべてを完成させる場所ではない。
師が守ってきた火を、弟子が別の形で燃え上がらせる場所である。
ヒントンの工房も、そうだった。
ヒントンは、長年の信念と直観を持っていた。
クリジェフスキーは、それを現実に走らせる実装力を持っていた。
サツケヴァーは、その先に広がる理論的地平を見ていた。
そして、ある時、三人は、決断した。
『来たる二〇一二年のImageNet Challengeに、出よう』
当時、画像認識の主流は手作業で設計した特徴量だった。人間が画像のどこを見るべきかを考え、輪郭、局所特徴、勾配、模様を取り出し、それを分類器に渡す。職人芸の世界である。
神経網が勝つと本気で信じる者は、まだ多くなかった。
だが、ヒントン工房の三人は、違ったものを見ていた。
特徴は、人間が削り出すものではない。
機械自身が、データから学び取るものだ。
その信念を、彼らは一つのネットワーク、深層学習に込めた。

五 ディープラーニングの衝撃

二〇一二年秋。ImageNet Challengeの結果が発表された。
ヒントン工房から提出されたシステムは、当時『SuperVision』と呼ばれていた。後に人々は、それをAlexNetと呼ぶようになる。主な実装者、アレックス・クリジェフスキーの名にちなんで。AlexNetは、八層の畳み込みニューラルネットワークだった。五つの畳み込み層と三つの全結合層。約六千万個のパラメータ。二枚のNVIDIA製GPUで訓練された。
そこには、いくつもの工夫があった。ReLU。
ドロップアウト。
データ拡張。GPUによる高速計算。
だが、最も重要だったのは、それらが一つの実用的な巨大システムとして組み上げられたことだった。
結果は、衝撃的だった。AlexNetのトップ五誤差率は、約一五・三パーセント。
二位のシステムは、約二六・二パーセント。
差は、一〇パーセントを超えていた。
研究の競技では、一パーセントの差でも大きい。数パーセントなら決定的である。だが、一〇パーセントを超える差は、もはや同じ延長線上の改良ではなかった。
それは、時代の断層だった。
人々は悟った。
これは偶然ではない。
これは小手先の調整ではない。
深層学習は、本物である。
長いあいだ時代遅れと呼ばれてきた神経網が、画像認識の頂点に立った。
クリジェフスキーは、いつものように淡々としていたという。
サツケヴァーは、この勝利の先にある大きな変化を感じていた。
ヒントンは、静かだった。
彼は、勝ち誇る必要がなかった。
冬を越した者にとって、春は大声で宣言するものではない。
ただ、雪が融けたことを知るだけでよい。
ヒントンの勝利は、静かな勝利だった。

六 ルネサンスの開幕

AlexNetの勝利を境に、世界は、文字通り、一夜にして変わった。
その『一夜』は、誇張ではなかった。
二〇一二年十月——AlexNetの結果が世に出てから、わずか数か月で、業界の言語そのものが変わった。Google、Microsoft、Facebook——大企業が、一斉に神経網研究者を、高給で引き抜き始めた。前年まで、ほとんど誰も関心を持たなかった研究分野が、突然、世界最先端の最重要技術となったのである。
ヒントン自身も、決断を下した。
二〇一二年末、彼は、自身が立ち上げた小さな会社『DNNresearch』を、競売にかけた。
会社といっても、そこに巨大な設備があったわけではない。製品が並んでいたわけでもない。あったのは、ヒントン、クリジェフスキー、サツケヴァーという三人の頭脳と、彼らが持つ技術だった。Google、Microsoft、Baidu、DeepMind——四社が入札した。
翌二〇一三年、GoogleはDNNresearchを買収する。金額は約四千四百万ドルと報じられた。
三人の頭脳に、世界の巨大企業がそれだけの価値を見たのである。
ヒントンはGoogleへ加わった。クリジェフスキーもまた、しばらくその流れの中に身を置いた。サツケヴァーは後に別の道を歩む。二〇一五年、OpenAIの創立に加わり、やがてGPT系列の技術的中核を担うことになる。
だが、それはまだ先の話である。
二〇一二年から二〇一三年にかけて起きたことは、単なる技術流行ではなかった。
それは、ルネサンスの開幕だった。
フィレンツェの工房で磨かれた技法が、やがてローマへ、ヴェネツィアへ、フランドルへ広がったように、ヒントン工房の火は世界中の研究室へ移っていった。
もはや神経網は、片隅の異端ではなかった。
深層学習は、AI研究の中心語になった。
長い中世が、終わったのである。

終 工房から世界へ

ヒントン工房から、世界へと、弟子たちは広がっていった。
クリジェフスキーは、しばらくGoogleに在籍した後、AIから距離を置く時期もあった。
サツケヴァーは、OpenAIへ。後にその主任研究員として、GPT系列のモデル開発の中核を担うことになる。さらに後年、OpenAIを去り、別の研究組織を立ち上げる——だが、これはずっと先の話である。
ヒントン自身は、Googleで研究を続け、二〇一八年、ヤン・ルカン、ヨシュア・ベンジオと共に、計算機科学の最高峰であるチューリング賞を受賞する。
冬を生き延びた三人が、ついに時代の中心へ迎えられたのである。
さらに、ベンジオの系譜からはイアン・グッドフェローが現れる。
二〇一四年、彼は敵対的生成ネットワーク、GANを発表した。
二つのネットワークを競わせることで、データに似た新しいものを生成する。
GANは、深層学習による画像生成を大きく前進させた。
その後、別の系譜から拡散モデルが台頭し、Stable Diffusionのような画像生成AIへつながっていく。
一つの発明が、そのまま直線的に次の発明を生むわけではない。
工房の火は、枝分かれし、別の炉で別の燃え方をする。
弟子が師となり、また次の弟子へ渡す。
そうして技術は、人物の系譜として広がっていく。
ヴェロッキオの工房からレオナルドが出たように。
蘭学者の小部屋から、近代の知が広がったように。
ヒントンの工房から、二十一世紀のAIが広がっていった。
ローゼンブラットの蝋の翼は、一度、海へ落ちた。
ヒントンたちは、その翼を拾い、別の素材で作り直した。
GPUとデータと学習アルゴリズムでできた、新しい翼である。
その翼は、二〇一二年、ついに飛んだ。
そして、その飛翔は止まらなかった。
炎は、まず画像を照らした。
次に、音声を照らした。
やがて、言葉を照らす。
そして二〇一六年、囲碁の盤上で、世界は再び機械の一手に息を呑むことになる。
その一手は、のちに『神の一手』と呼ばれる。
第六章 了

第七章 二つの神の一手

AlphaGo対李世乭(二〇一六)

一 四千年の盤

囲碁は、古い競技である。
中国の伝説によれば、紀元前二三〇〇年頃、伝説の帝堯が、出来の悪い息子・丹朱を教育するために、囲碁を発明したと伝えられる。それから数えれば、四千年以上の歴史を持つ。
盤は、十九路。縦横十九本の線が交差した、三六一の交点。 石は、二色。黒と白。 ルールは、極めて単純である。互いに石を置き合い、相手の石を囲んで取り、より多くの陣地を占めた者が勝つ。
それだけの規則で、宇宙の原子数を遥かに上回る、十の百七十乗もの局面が、可能となる。チェスの局面数がおよそ十の四十数乗とされるのに比べても、桁違いの深遠さである。
中国では、囲碁は『琴棋書画』の一つに数えられた。琴を弾き、碁を打ち、書をしたため、絵を描く。それは単なる遊戯ではなく、士大夫の教養であった。
日本では、徳川幕府が『本因坊』『井上』『安井』『林』の四家を保護し、将軍の前で『御城碁』なる対局を行わせた。本因坊算砂、本因坊道策、本因坊秀策——歴代の名人たちは、半ば神格化された。
朝鮮半島でも、囲碁は深く根づいた。二十世紀後半以降、韓国は世界最強の囲碁国家の一つとなり、数多くの天才棋士を生んだ。
二十世紀末、機械はチェスを攻略した。
だが囲碁は違う、と人々は考えていた。
チェスは、駒の動きが明確で、局面の評価も比較的定式化しやすい。もちろん深いゲームである。だが、囲碁は可能な局面の数が、桁違いに多い。盤面の評価が、極めて曖昧である。『良い形』と『悪い形』を、機械に教えるのは、ほとんど不可能とされた。
これらをどう数式にするのか。どう機械に教えるのか。
よい手は、しばしば言葉にならない。名人は盤面を一目見て、『ここが急所だ』と感じる。その感覚こそ、機械には届かない最後の領域だと思われていた。
囲碁は、人間が機械から守り抜く最後の砦であった。
多くの専門家は言った。
機械がトップ棋士を破るには、まだ十年、あるいは二十年はかかるだろう。
二〇一六年三月。
その二十年は、まだ来ていないはずだった。

二 デミス・ハサビス

この物語のもう一人の主役は、英国人デミス・ハサビスである。
一九七六年、ロンドン生まれ。父はギリシャ系キプロス人、母はシンガポール華人。幼い頃から、彼は際立った才能を示した。
四歳でチェスを覚え、少年時代には英国チェス界の神童として知られるようになる。十三歳のころには、すでに世界有数のジュニア棋士の一人だった。
だが、彼はチェスの盤上だけに留まらなかった。
十代半ばで、彼はゲーム開発の世界に入った。十代後半に開発に関わった『テーマパーク(Theme Park)』——遊園地を経営するシミュレーションゲームは、欧州中でベストセラーとなった。
二十代には、自らゲーム会社エリクシール・スタジオを設立する。
そして三十代、彼は一度、ゲームの世界を離れた。
向かった先は、神経科学である。
ロンドン大学ユニバーシティ・カレッジ。そこで彼は、人間の記憶、とりわけ海馬が記憶や想像にどう関わるかを研究した。
なぜゲーム開発者が、脳を研究するのか。
ハサビスの答えは、一貫していた。
『私は、いつの日か、機械に知能を宿らせたいと考えていた。そのためには、まず、人間の脳が、どう知能を生み出しているのかを、深く理解せねばならぬと考えた』
知能を作るためには、まず知能を知らなければならない。
十五世紀のレオナルド・ダ・ヴィンチは、絵を描くために人体を解剖した。馬を描くために骨格を調べ、飛ぶ機械を考えるために鳥の翼を観察した。美しい表面を描くために、その内側の構造を見ようとした。
ハサビスもまた、人工知能を作るために、人間の知能の内側へ入っていった。
二〇〇九年、博士号取得。
二〇一〇年、彼は仲間とともにロンドンで新しい研究所を立ち上げる。DeepMind。
彼らが掲げた野心は大きかった。
知能を解明し、その知能によって世界の難問を解く。
それは、ゲーム会社でも、普通の大学研究室でもなかった。
人工知能のための、現代の工房であり、結社であった。

三 直観、経験、読み

DeepMindは、すぐに世界を驚かせる。
二〇一三年、彼らはAtari 2600のゲームを機械に学ばせる研究を発表した。
最初の論文で試されたのは、七つのゲームだった。
機械に与えられるのは、画面の生のピクセルと、取ることのできる行動、そしてゲームから返される報酬である。
ルールブックを読ませるわけではない。
『ブロックとは何か』『ボールとは何か』と人間が教えるわけでもない。
機械は、画面を見て、行動し、得点という報酬を受け取り、試行錯誤を繰り返す。
その結果、七つのうちいくつかでは従来法を大きく上回り、三つのゲームでは人間の熟練者を超える成績を示した。
深層神経網と強化学習の結合である。
二〇一四年、GoogleはDeepMindを買収した。若い研究所は、巨大な計算資源を手に入れた。
次の標的は、囲碁だった。
DeepMindが作ったAlphaGoは、一つの技術だけでできていたわけではない。そこには三つの力が組み合わされていた。
第一に、深層神経網——盤面のパターンから、有望な手や局面の価値を推定する力。
第二に、強化学習——自己対局を重ね、戦略を磨き上げる仕掛け。
第三に、モンテカルロ木探索——可能性のある手を、効率的に読み解く探索。
これら三つを、見事に融合させた。
深層神経網が直観に似た役割を担う。
強化学習が経験を積む。
木探索が読みを深める。
——直観と、経験と、読み。
人間の囲碁棋士の言葉に引き寄せて表現するなら、そういう三つの力である。
ハサビスらは、この機械に、AlphaGoという名を付けた。

四 前哨戦

二〇一五年十月、ロンドン。DeepMindの研究所に、一人の棋士が招かれた。
樊麾。中国出身、フランス在住。ヨーロッパ囲碁王者。プロの棋士として、二段の段位を持つ。
公衆には、まだAlphaGoの存在は、明らかにされていなかった。これは、極秘の前哨戦であった。
五局が打たれた。
結果は、AlphaGoの五戦全勝。
プロ棋士が、互先(ハンデキャップのない対局)でコンピュータに敗れた。囲碁の歴史における重大な転換だった。
だが、世界の反応はまだ半信半疑だった。
樊麾は欧州の王者である。だが、真の頂点は韓国と中国にいる。世界最強級の棋士なら、まだ負けない。そう考える者は多かった。DeepMindは、次の相手を探した。
選ばれたのは、韓国の李世乭九段である。
李世乭は、単に強いだけの棋士ではなかった。鋭く、激しく、常識外の手を好む。均整の取れた完璧さというより、盤面を破壊して新しい秩序を作り出す棋士だった。AlphaGoが人間の直観を超える存在なら、その相手にふさわしいのは、人間の側で最も創造的な棋士の一人であるべきだった。
ソウルでの五番勝負が決まった。
世界の多くは、まだ人間の勝利を信じていた。

五 第二局・第三十七手

二〇一六年三月九日、ソウル。
フォーシーズンズホテルに、対局室が設けられた。
盤。
石。
時計。
二つの椅子。
対局者の一方は、李世乭九段。当時、三十三歳。韓国を代表する世界最強級の棋士。十年以上、世界のトップ集団の中で戦い続けてきた、現代囲碁界の頂点の一人。
もう一方は——AlphaGo。
機械は、自分で石を盤に置くことはできない。代わりに石を置くのは、開発に関わった台湾系の若き研究者、黄士傑(アジャ・ファン)博士。彼は、画面に表示されるAlphaGoの判断を読み、その手を、盤に置く役割を担っていた。
第一局——AlphaGoが、勝利した。世界が、息を呑んだ。
そして——第二局。
三月十日、午後。AlphaGoの第三十七手。
それは、人間なら、絶対に打たない位置に、置かれた石であった。
その石が置かれた瞬間、観戦していたプロ棋士たちは困惑した。
それは、当時の人間の感覚ではほとんど選ばれない場所だった。五線への肩ツキ。序盤でそこへ打つのは、形が悪い、早すぎる、意味が薄い——そう見えた。
会場の一隅で、樊麾——前年に同じAlphaGoに敗れた、あの樊麾——は、しばし言葉を失ったという。彼は後に、こう語っている。
『これは、人間の手ではない。こんな手を打つ人間を、見たことがない。なんと、美しい』
しかし、局面が、進んでいくにつれて——様相が変わった。
その『悪い形』に見えた一手が、後の展開で、実は極めて深い意味を持つことが、徐々に明らかになってきた。盤面の中央部での影響力が、その一手によって、確保されていた。
李世乭は、終局後、こう語った。
『私は、コンピュータが創造的な手を打つとは、思っていなかった。だが、第三十七手——あれは、明らかに、創造的だった。美しかった、と認めるしかない』
これは、人類の囲碁の常識を超えた、もう一つの『神の一手』だった。
何千年もの囲碁の伝統の中で、誰も打ったことのない、しかし正しい手——AlphaGoは、人類が四千年かけて築いた囲碁の智慧の、まだ知らぬ領域を、発見していたのである。
第二局も、AlphaGoが勝った。
第三局も、AlphaGoが勝った。
三連敗。
人類の最強の代表が、機械に三度、連続して敗れた。
夜のホテルの記者会見場で、李世乭は深く頭を下げた。
『多くの方々の期待に応えられず、申し訳なく思う。これほどの重圧を感じたことは、かつてなかった』
世界は、この三連敗を、人類の敗北として受け取ろうとしていた。だが、当の李世乭は、その物語を退けている。これは李世乭の敗北であって、人類の敗北ではない。示されたのは私の弱さであって、人類の弱さではない。のちに彼は、そう言い切った。

六 第四局・第七十八手

三月十三日、第四局。
五番勝負としては、すでに決着がついていた。AlphaGoの三勝。残り二局で、李世乭が勝ち越すことはできない。
それでも彼は、盤の前に座った。
人間が、なお一手を返せるかどうか。
少なくとも、世界はそのような物語として、この一局を見ていた。
序盤から中盤にかけて、AlphaGoは優勢に見えた。観戦者の多くも、また同じ結末を予想し始めていた。
そして、第七十八手。
李世乭は、盤面の中央に、ある妙手を打ち込んだ。
DeepMindは後に、この手を、通常なら一万回に一度ほどしか選ばれないような手として紹介した。
極めて意外な一手だった。
その石が置かれた後、局面は揺れた。
もちろん、機械に感情はない。動揺も恐怖もない。
だが、盤上に現れた手順は、まるで機械が平衡を失ったかのように見えた。
AlphaGoはその後、形勢を損ねていく。
そして、ついに投了。
AlphaGoが負けた。
李世乭が、二〇一六年の五番勝負で、AlphaGoから唯一の一勝を挙げた瞬間である。
世界は、その第七十八手を『神の一手』と呼んだ。
これが、第二の神の一手である。
第一の神の一手は、機械の側から来た。
AlphaGoの第三十七手。
人類の常識の外から現れ、人間の美意識を揺さぶった手。
第二の神の一手は、人間の側から来た。
李世乭の第七十八手。
機械の予測の外から現れ、ただ一度、この五番勝負の巨大なシステムを沈黙させた手。
この二つの手が、同じ五番勝負の中で生まれた。
それは、あまりにも美しい対称だった。
機械は、人間の知らなかった囲碁を見せた。
人間は、機械の知らなかった一手を見せた。
神は、どちらか一方の側にいたのではない。
盤の上の可能性そのものに、宿っていたのである。

終 勝てない存在のあとで

最終結果は、四勝一敗。AlphaGoの完勝だった。
しかし、世界の記憶に最も鮮やかに残ったのは、AlphaGoの四勝ではなく、李世乭の一勝だった。なぜなら、そこには人間の尊厳が、敗北の中で光る瞬間があったからである。
対局シリーズの後、ハサビスは、李世乭に深く感謝を述べた。
『私たちは、機械が人間を倒すために、AlphaGoを作ったのではない。人間が、新しい何かに気づくために作った。あなたの第七十八手は、私たちの機械にも、新しいことを教えてくれた』
李世乭は、しばらく沈黙していたが、こう答えた。
『私は、機械に負けたとは、思いたくない。私は、新しい囲碁に出会った、と思いたい』
二〇一七年、AlphaGoはさらに進化し、中国の世界一位、柯潔(カ・ケツ)九段を破った。柯潔は、対局後、涙を流したと伝えられる。人間の頂点に立つ棋士が、もはや届かない高さを前にした涙だった。
その後、AlphaGoは現役を退いた。だが、その子孫たちは残った。囲碁AIは急速に普及し、プロ棋士もアマチュアも、AIの示す候補手を見ながら研究する時代へ入っていく。
二〇一九年、李世乭は引退を表明した。
引退の理由を問われ、彼はこう答えた。
『AIの登場により、私は、自分が一位を目指す意味を、見失った。たとえ私が世界一位になっても、その上に、勝てない存在がいる。それを知ってから、囲碁を続ける動機が、薄らいでしまった』
そして、それは——人類全体が、これからAIと共に生きていく中で、繰り返し問われることになる問いの、最初の声であったのかもしれない。
『機械の方が優れた存在となった時、人間は、何のために、その道を歩むのか』
その問いに、まだ、答えは出ていない。
李世乭は、引退後、ソウルに残った。時に、子供たちに囲碁を教えた。
『人間が打つ囲碁には、人間の打つ囲碁にしかない何かがある』と、彼は、子供たちに語ったという。
その『何か』が、何であるか——彼自身も、まだ、はっきりとは、言葉にできなかった。
だが、それを、子供たちに伝えていきたい、と彼は願っていたのである。
四千年の盤は、まだ、置き続けられている。
第七章 了

第八章 八人の使徒

トランスフォーマー来航(二〇一七)

一 二〇一七年のグーグル・ブレイン

二〇一七年、カリフォルニア州マウンテンビュー。Google本社の広大な敷地には、ひとつの知的な工房があった。
その名を、Google Brainという。
二〇一一年、アンドリュー・ン、ジェフ・ディーンらによって始められたこの研究部門は、AlexNet以後の深層学習の熱を、企業の中心へと運び込んだ場所であった。Googleはそこに莫大な資金と計算資源を注ぎ、世界中から神経網研究の俊才を集めていた。
彼らの戦場のひとつが、自然言語処理であった。
人間の言葉を、機械に理解させ、翻訳させ、生成させる。
一見すれば単純な願いである。だが、言葉ほど手強いものはない。
チェスには盤面がある。囲碁には十九路盤がある。
しかし言葉には、明確な盤がない。
そこには、文法があり、歴史があり、文化があり、皮肉があり、沈黙があり、言外の意味がある。
一語の背後に、千年の記憶が潜むことすらある。
その流れを、機械に扱わせる。Google Brainの研究者たちは、日々、その難題に向き合っていた。
そして二〇一七年の春、その研究室の周辺で、八人の研究者がひとつの構想へと近づいていた。
アシシュ・ヴァスワニ。インド出身の若き計算機科学者。
ノーム・シェイザー。Googleの古参研究員。長く言語モデルに取り組んできた人物。
ニキ・パルマー。インド出身の研究者。
ヤコブ・ウシュコライト。ドイツ生まれの機械翻訳の専門家。
リオン・ジョーンズ。ウェールズ出身の英国人。
エイダン・ゴメス。カナダ、トロント大学から来たインターン。当時まだ二十歳そこそこの学部生であった。
ウカシュ・カイザー。ポーランド出身の研究者。
イリヤ・ポロスーキン。ウクライナ出身の研究者。
国籍も違う。年齢も違う。経歴も違う。
共通していたのは、ただ一つ——彼らがGoogle Brainで、自然言語処理に取り組んでいた、ということだけであった。
そして、彼らの中に、ある革新的な発想が、芽生えつつあった。
それは、自然言語処理を、根本から変える発想であった。

二 RNNの壁

彼らの発想を理解するには、当時の自然言語処理が、どこで行き詰まっていたのかを見なければならない。
二〇一七年当時、自然言語処理の主流は、リカレント神経網、すなわちRNNであった。RNNは、文章を、一語ずつ、順番に処理する。最初の単語を読み、内部状態を更新する。次の単語を読み、また内部状態を更新する。これを、文の最後まで繰り返す。
これは、人間が文章を読む過程に、似ていた。最初から最後まで、順番に。
だが、RNNには、二つの致命的な弱点があった。
第一に、遅い。
文章を順番に処理するため、並列計算ができない。一つの単語の処理が終わるまで、次の単語の処理は始められない。長い文章になれば、処理時間は、文章の長さに比例して、伸びていく。
第二に、忘れる。
文章が長くなるにつれて、最初の方の単語の情報が、内部状態の中で薄れていく。例えば、長い物語の最後で『彼』という代名詞が出てきた時、それが物語の最初に登場した誰を指すのか——RNNは、その記憶を失っていく。RNNは、その記憶を保ち続けることが苦手であった。
一九九〇年代には、LSTMという改良版が発明されていた。
これは記憶を長く保つための門を備えた、巧妙な構造であった。
のちにはGRUのような簡略化された変種も生まれた。
それでもなお、根本は変わらなかった。
順番に読む。
順番に処理する。
前が終わらなければ、次へ行けない。
この逐次性こそが、自然言語処理の壁であった。
短い文ならばよい。
だが、記事全体を理解する。本一冊を理解する。長い対話の流れを保つ。
そうした仕事を考えると、RNNの足取りは、どうしても重くなった。
八人の研究者たちは、この壁を見ていた。
そして、ある時、発想の向きを変える。
文章は、本当に一語ずつ読まなければならないのか。

三 Attentionという発想

八人の試みを、一文に縮めてしまえば、こう言える。
『文章を、順番に処理することだけに頼るのを、やめてみたらどうだろう』
もちろん、実際の研究は、そんな一言から突然生まれたわけではない。
Attentionという考え方は、それ以前から機械翻訳などで使われていた。
文章のある部分を処理するとき、入力のどこを強く参照すべきかを学ぶ仕組みである。
八人の発想が急進的だったのは、それを補助装置のままにしなかったことだった。
Attentionを、中心構造にする。
RNNを外す。
畳み込みも外す。
Self-Attentionによって、系列の中の各要素が、他の要素との関係を直接計算する。
たとえば、『私は猫が好きです』という文章がある。
『好き』という語を理解するには、『誰が』好きなのか、『何を』好きなのかという関係が重要になる。
Self-Attentionは、そうした関係を重みとして計算する。
この仕組みには、大きな利点があった。
第一に、並列計算がしやすい。
RNNのように、前の時刻の計算が終わるまで次を待つ必要がない。
GPUやTPUの力を、はるかに活かしやすくなる。
第二に、遠い要素同士を直接結びつけられる。
文章の冒頭と末尾が、何十段もの記憶の受け渡しを経ずに関係を持てる。
もちろん、順番を完全に捨てるわけにはいかない。
『猫が私を好き』と『私が猫を好き』は、同じ単語でも意味が違う。
そのため、位置の情報を別に加える。
単語そのものと、単語同士の関係と、位置。
それらを組み合わせる。
こうして彼らは、新しい構造を組み立てた。
名前は、Transformer。
変換するもの。
控えめな名前である。
だが、その構造は、自然言語処理の歴史を根本から変えることになる。

四 ビートルズの残響

論文の題名は、『Attention Is All You Need』。
必要なのは、Attentionだけ。
RNNも、畳み込みも使わず、Attentionを中心に組み立てる——という論文の大胆さを、そのまま表した題名である。
この題は、締切を目前にした慌ただしい会話の中で生まれたと、のちに当人たちが明かしている。RNNも畳み込みも捨て、Attentionただ一つに賭ける。その大胆さを言い当てる短い言葉を、彼らは探していた。
そのとき、命名者リオン・ジョーンズの頭に浮かんだのは、半世紀前の歌だったという。
一九六七年。
ビートルズは『All You Need Is Love』を歌った。
必要なのは愛だけ。
それから、ちょうど半世紀。
二〇一七年。
必要なのは、Attentionだけ。
題名の由来は、のちにジョーンズ自身が、あの歌にちなんだと明かしている。
ここから先は、その事実の上に置く、本書の読みである。
愛から注意へ。
歌から論文へ。
人間の世界を変えた合唱と、機械の言語を変えた数式が、同じ構文の中で半世紀を隔てて響き合う。
奇妙な符合である。
『注意』という語は、人間に属する言葉でもある。
誰かに注意を向ける。
何かを気にかける。
関係を結ぶ。
その言葉が、機械の知性を支える中心概念になった。
『Attention Is All You Need』。
題名だけで、すでに宣言だった。

五 論文公開

二〇一七年六月十二日。arXiv——世界中の研究者が、論文を事前公開する、無料の文書サーバ。物理学者から始まり、今では計算機科学者も、頻繁に利用する場所。
そこに、八人の論文が、公開された。Ashish Vaswani。Noam Shazeer。Niki Parmar。Jakob Uszkoreit。Llion Jones。Aidan N. Gomez。Łukasz Kaiser。Illia Polosukhin。
題名は、『Attention Is All You Need』。
論文は、わずか十数ページであった。簡潔な構造図、簡潔な数式、簡潔な実験結果。
公開直後の反応は、極めて穏やかであった。
業界の主要な研究者の何人かは、この論文を読んだ。『興味深い実験だ』『いくつかの良いアイデアがある』と評価する者もいた。だが、誰一人として、これがAI研究の地殻変動になるとは、思っていなかった。
歴史を変える論文は、しばしばそうして現れる。
一六八七年、アイザック・ニュートンは『自然哲学の数学的諸原理』を世に出した。
ラテン語で、『Philosophiæ Naturalis Principia Mathematica』。
のちに『プリンキピア』と呼ばれる書物である。
その本は、万有引力と運動法則を数学の言葉で記述した。
しかし、刊行当時、それを本当に理解できた者は限られていたと言われる。
英国王立協会の会員でさえ、その全貌をすぐに掴んだ者は少なかった。
それでも『プリンキピア』は、時間をかけて世界を変えた。
天体の運行も、砲弾の軌道も、潮の満ち引きも、すべて同じ数学の言葉で語れることを示したからである。
二〇一七年六月のarXivに置かれた『Attention Is All You Need』も、似た運命をたどる。
最初は、静かな論文であった。
だが、そこには、現代AIの力学が記されていた。
二〇一八年——Googleが『BERT』を、OpenAIが『GPT-1』を発表する。両者とも、トランスフォーマー構造を、根幹に据えていた。
二〇一九年——『GPT-2』。 二〇二〇年——『GPT-3』。
そして、二〇二二年——ChatGPTが、世界を一変させる。
世界中の人々が、機械と自然な言葉で対話する時代が訪れる。
その建築物の基礎には、二〇一七年六月十二日の、あの十数ページの論文があった。
ニュートンの『プリンキピア』が近代物理学の基礎となったように、
『Attention Is All You Need』は、現代AIの基礎文献となった。
そこに書かれていたのは、単なる翻訳モデルではない。
機械が言葉を扱うための、新しい重力法則であった。

六 ディアドコイの分散

では、その論文を書いた八人は、その後どうなったのか。
ここに、もう一つの歴史的比喩が現れる。
八人は、やがてGoogleを離れ、それぞれ別の場所へ散っていった。
筆頭著者アシシュ・ヴァスワニは、ニキ・パルマーらとともに、Adept AIの創業に関わり、のちにEssential AIを立ち上げた。
ノーム・シェイザーは、ダニエル・デ・フレイタスとともにCharacter.AIを創業した。対話型AIの巨大な潮流を生み出す会社である。
エイダン・ゴメスはカナダへ戻り、Cohereを共同創業した。企業向け大規模言語モデルの有力企業として、急速に存在感を増していく。
ヤコブ・ウシュコライトは、生物学とAIを結ぶ方向へ進み、Inceptiveを立ち上げた。mRNA医薬への応用を目指す、異色の結社である。
イリヤ・ポロスーキンは、ブロックチェーンの世界へ転じ、NEAR Protocolを共同創業した。
ウカシュ・カイザーは、OpenAIへ移った。
そしてリオン・ジョーンズは、東方へ向かう。
まるで、アレクサンドロス大王の死後に現れた、ディアドコイのようであった。
紀元前三二三年、アレクサンドロス大王が、三十二歳の若さでバビロンに病死した。彼が残した、地中海から中央アジアに広がる大帝国を、誰が継ぐのか。
その大将軍たちは、互いに戦い、和議を結び、また戦い——最終的に、帝国を分割した。プトレマイオスがエジプトを、セレウコスがアジアを、カッサンドロスがマケドニアを、リシマコスがトラキアを——それぞれ後継国家として継承した。
これら後継国家は、それぞれが独自に発展し、独自の文化を花開かせた。プトレマイオス朝のアレクサンドリア図書館。セレウコス朝のヘレニズム文化。彼らはアレクサンドロス大王の遺産を、各地で別の形に育てていった。
トランスフォーマーの八人にも、同じ構造があった。Google Brainという巨大な都から、彼らは散っていった。
それぞれが、新しい結社を築いた。Essential AI、Character.AI、Cohere、Inceptive、NEAR Protocol、OpenAI、Sakana AI。
それらは単なる会社名ではない。
トランスフォーマーという発明が、世界中で異なる形に変奏されていく、その拠点であった。
ある者は対話AIへ。
ある者は企業向け言語モデルへ。
ある者は医療へ。
ある者はブロックチェーンへ。
ある者は、さらに巨大な基盤モデルへ。
そしてある者は、東京へ。
ディアドコイがヘレニズム文化を各地へ広げたように、
八人の使徒は、トランスフォーマーの遺産を世界へ広げた。
論文は一つだった。
だが、その後に生まれた王国は、一つではなかった。

終 東京のリオン・ジョーンズ

八人の中で、もっとも興味深い末路を辿った者がいる。
リオン・ジョーンズ。
ウェールズに生まれ、バーミンガム大学で計算機科学を学んだ。Googleに入り、Google Brainで研究に携わり、『Attention Is All You Need』の共著者となった。
論文公開後もしばらくGoogleに残った彼は、二〇二三年、ひとつの決断を下す。
東京へ渡る。
東京で、彼は、もう一人の元Google研究員と組んだ。デイヴィッド・ハー。香港生まれのカナダ人。長年、Google Brain Tokyoで、進化計算と自然インスピレーションに基づくAI研究を続けていた人物であった。
二人は、元外交官の伊藤錬を加え、東京で新しい会社を立ち上げた。Sakana AI。
魚の群れ。
これが、二人が選んだ名前であった。
自然界では、無数の魚が、群れで泳ぐ。一匹一匹は、それほど賢くない。だが、群れ全体としては、極めて巧妙な動きをする。捕食者から逃れ、餌を見つけ、集団で最適な行動を取る。
二人は、これを、AI研究の理想として、掲げた。
巨大な単一のモデルではなく——多数の小さなモデルが、群れとして協力する形のAIを、彼らは追求した。
これは、当時の業界の主流——『より大きなモデルが、より賢い』という考え——への、静かな反論であった。
そして、それを、東京の地で——日本のAI業界の中で——追求するという、二人の選択。
かつて第五世代コンピュータの壮夢を掲げ、AIの冬も春も経験した日本の地に、トランスフォーマー論文の使徒の一人が根を下ろしたのである。
プトレマイオスはエジプトに渡り、アレクサンドリアに知の都を築いた。
セレウコスはバビロンを基点に、東方へ広がる王国を築いた。
リオン・ジョーンズは東京に渡り、自然に学ぶAIの拠点を築こうとした。
二〇一七年六月十二日、arXivに置かれた一篇の論文。その余波は、マウンテンビューから世界へ広がり、やがて東京の街角にも届いた。
『プリンキピア』が近代世界の力学を変えたように、『Attention Is All You Need』は、知能の力学を変えた。
そしてディアドコイが帝国の遺産を各地で花開かせたように、八人の使徒は、それぞれの地で新しいAIの王国を築いていった。
歴史は、ときに一篇の論文から始まる。
そしてその論文は、ときに海を渡る。
第八章 了

第九章 西へ

GPT-1からGPT-3への航海(二〇一八—二〇二〇)

一 出航前夜

二〇一八年初頭、サンフランシスコ。OpenAIの研究所は、まだ小さかった。研究者と技術者を合わせても、百名に満たぬ規模である。創設時の庇護者であったイーロン・マスクは、この年、理事会を離れていく。Googleが抱える巨大な計算資源と人材の厚みに比べれば、OpenAIは、まだ港に繋がれた一艘の帆船にすぎなかった。
それでも、その甲板には奇妙な熱があった。
彼らは、何かを掴みかけていた。
その『何か』が何であるのか、まだ誰にも名付けられなかった。だが、水平線の向こうに、見たことのない陸影があるような予感だけはあった。
前年、二〇一七年六月。Google Brainの八人が『Attention Is All You Need』を世に問うた。トランスフォーマーという新しい構造は、言語処理の海に投げ込まれた羅針盤であった。OpenAIの研究員たちは、それをすぐに手に取った。
その中に、アレック・ラドフォードという若い研究者がいた。
彼は、ある興味深い構想を、温めていた。
『トランスフォーマーを、大量のテキストで、ただ予測だけをさせる』
それだけである。
次に来る単語は何か。正確には、次に来るトークンは何か。
機械はそれを、来る日も来る日も予測する。間違えれば重みを直し、当たればその方向を強める。何億回、何十億回と繰り返す。
あまりに単純な作業であった。
だが、ラドフォードたちは考えた。
もし、この単純な予測を、十分に大きな規模で続けたらどうなるのか。
一四九二年、コロンブスは『西へ向かえば、いつかインドに着く』と信じ、大西洋へ船を出した。海図は粗く、距離の見積もりは誤っていた。それでも、彼は西へ向かった。
ラドフォードたちもまた、同じように西を見ていた。
十分に大きなモデル。
十分に多いテキスト。
十分に長い学習。
その先に、何があるのか。
誰にもわからなかった。
だからこそ、彼らは船を出した。

二 最初の船——GPT-1

二〇一八年六月。OpenAIから一篇の論文が発表された。
題名は『生成的事前学習による言語理解の改善』。
著者は、アレック・ラドフォード、カーティク・ナラシムハン、ティム・サリマンス、そしてイリヤ・サツケヴァー。
サツケヴァーは、かつてヒントン工房にいた弟子であり、AlexNetの共著者でもあった。トロントから深層学習の春を運び、いまはOpenAIの技術的中核を担っていた。
彼らが世に出したモデルの名は、GPT-1。GPTとは、Generative Pre-trained Transformer——生成的事前学習トランスフォーマーの略である。
パラメータ数は、一億一千七百万。
当時の言語モデルとしては、大きい部類ではあった。だが、巨大というほどではなかった。GPT-1の核心は、構造そのものの新しさではない。トランスフォーマーは、すでにGoogleが発表していた。OpenAIが見出したのは、その船をどのように使えば外海へ出られるか、という航法であった。
航法は、二段階だった。
第一段階は、事前学習(Pre-training)。BooksCorpusと呼ばれる、未出版書籍を中心とした大規模なテキスト集合を機械に読ませる。およそ七千冊、八億語規模の文章である。機械はそこで、ただ次の語を予測し続ける。
人間が明示的に文法を教えるわけではない。
『これは主語である』『これは比喩である』『これは皮肉である』と、誰かが札を貼るわけでもない。
ただ文章を読み、次を当てる。
しかし、その単純な訓練の中で、機械は少しずつ、人間の言葉の癖を掴んでいく。単語同士の距離、文法の型、物語の流れ、問いと答えの形。言語の海流を、暗黙のうちに身につけていく。
第二段階は、微調整(Fine-tuning)。
事前学習を終えたモデルに、今度は特定の課題を与える。文章分類、質問応答、含意判定。少量の教師データを用いて、目的に合わせて舵を調整するのである。
ラドフォードたちの仮説は、こうだった。
『第一段階で、機械は言語の一般的な知識を獲得する。第二段階で、それを特定の課題に応用する。一つの基盤モデルから、無数の応用が、生まれるはずだ』
これが、現代のあらゆる大規模言語モデルの、最も基本的な作り方になった。OpenAIの最初の船は、地中海を、静かに、出航したのである。

三 BERTの平行航海

GPT-1の発表から、わずか四か月後。
二〇一八年十月、Googleが別の船を進水させた。
名を、BERTという。
ジェイコブ・デブリンらによるこのモデルは、GPT-1とは異なる思想で作られていた。GPT-1は、左から右へ読む。
すでに読んだ言葉をもとに、次の言葉を予測する。矢印は一方向に伸びていく。BERTは違った。
文章の一部を隠し、その隠された語を前後の文脈から当てる。
文章全体を見渡し、右からも左からも意味を汲み取る。双方向のモデルであった。
たとえば、『彼は銀行に行き、口座を開いた』という文がある。
『銀行』を理解するには、その前後を見なければならない。川の土手ではなく、金融機関であることは、『口座』という後ろの語が教えてくれる。BERTは、そのような前後関係を巧みに扱った。
この方法は、長い文章を理解するのに適していた。GPT-1が一億一千七百万パラメータであったのに対し、BERT-Largeは三億四千万パラメータ。およそ三倍の規模である。各種の言語理解課題で、BERTは当時の最高性能を次々と塗り替えた。
研究界の注目は、急速にBERTへ集まった。
業界の注目は、急速にBERTに集まった。GPT-1は、その影に隠れる形となった。
だが、OpenAIの研究員たちは、そこで引き返さなかった。彼らは別のものを見ていた。BERTは、文章を理解することに強かった。GPTは、文章を生成することに向いていた。BERTは、地図を広げて全体を読む船であった。GPTは、水平線の先へ文章を伸ばしていく船であった。
その違いは、当時、まだ小さな差に見えた。
だが後に、それは決定的な意味を持つ。OpenAIが目指していたのは、特定の課題での最高性能ではなかった。彼らが目指していたのは、もっと根本的な何か——『一つのモデルが、多様な課題を、ほぼ何の微調整もなく、こなす』という、夢のような能力であった。
そのためには、もっと大きな船が、必要であった。
地中海ではない。大西洋に出るための船が。

四 GPT-2と『危険すぎる』騒動

二〇一九年二月。OpenAIは、再び論文を発表した。
題名は『言語モデルは教師なし多課題学習者である』。
発表されたモデルは、GPT-2。
パラメータ数は、十五億。GPT-1の十倍を超える規模であった。
学習に用いられたのは、WebTextと呼ばれる大規模なウェブ由来のテキスト集合である。人間がリンクとして共有したページを集め、そこから文章を抽出した。海は、書物の内海から、インターネットの外洋へ広がっていた。GPT-2は、これまでよりはるかに自然な文章を書いた。
冒頭を与えれば、その続きを数段落にわたって書き継ぐ。ニュース記事のようにも、小説のようにも、論説のようにも振る舞う。内容はしばしば怪しく、事実でないことも平然と混じった。だが、その文体は滑らかだった。
そしてOpenAIは、論文と同時に、極めて異例の発表を行った。
『このモデルは、虚偽のニュースの量産、なりすまし文章の生成、悪意あるスパムなど、社会に害をもたらす用途に転用される危険が高い。よって、完全な形での公開は、見送る』
業界が、騒然となった。
『過剰反応である』という批判が、噴出した。『研究の進歩を、自分たちの判断で止めるのか』と。
逆に、『責任ある慎重さである』と評価する声も、あった。『強力な技術を、慎重に世に出すという姿勢は、評価すべきだ』と。
賛否は、激しく分かれた。
そして、この騒動の背後で、OpenAIの内部にも、亀裂が走り始めていた。
より安全性を重んじる者。
より速い展開を求める者。
公開による進歩を信じる者。
制御なき拡散を恐れる者。
その緊張の中に、ダリオ・アモデイの姿があった。当時OpenAIの研究担当副社長であり、のちに妹ダニエラ・アモデイらとともにAnthropicを創業する人物である。
だが、それはまだ先の話であり、二〇一九年の時点で、問いはひとつの形を取り始めただけである。GPT-2を世に出すか、出さないか——という議論は、後に、AI業界全体を分裂させる、より大きな問いの、最初の声であった。
『強力な技術を、誰が、どのように、世に出すべきか』
この問いは、まだ、答えが出ていない。

五 航海士カプランの法則

二〇二〇年一月。
OpenAIから、一篇の論文が出た。
題名は『神経言語モデルのスケーリング則』。
筆頭著者のジャレッド・カプランは、もともと物理学者であった。
彼らが問うたのは、単純で巨大な問題だった。
モデルを大きくしたら、性能はどう変わるのか。
データを増やしたら、何が起こるのか。
計算量を注ぎ込めば、どこまで良くなるのか。
それまで研究者たちは、経験的に知っていた。
大きくすれば、しばしば良くなる。
データを増やせば、しばしば強くなる。
だが、『しばしば』は海図ではない。
カプランたちは、規模の異なる多数の言語モデルを調べ、損失の下がり方を測った。
モデルサイズ。
データ量。
計算量。
そこには、広い範囲で驚くほど規則的な、べき乗則に近い関係が見えた。
規模を増やしたとき、性能がどのように改善するかを、ある程度予測できる。
大航海時代の航海士たちが、風と海流の規則性を知ったとき、航海は少しだけ賭けではなくなった。
スケーリング則もまた、巨大な言語モデルの海に、数式の海図を与えた。
ただし、ここで因果関係を単純にしてはいけない。
スケーリング則の論文が出てから、OpenAIが初めてGPT-3を思いついたわけではない。
海図が描かれるのと、巨大な船が建造されるのは、ほとんど同じ時期に進んでいた。
大規模化の実践が、法則を見つけさせた。
法則がまた、大規模化への確信を強めた。
理論と実践は、同じ港で互いを押し出していたのである。
そして数か月後、その港から、一七五〇億の船が出る。

六 一七五〇億——新大陸

二〇二〇年五月。
OpenAIから、一篇の巨大な論文が発表された。
題名は『言語モデルは少例学習者である』。
筆頭著者はトム・ブラウン。著者は三十一名に及んだ。
発表されたモデルの名は、GPT-3。
パラメータ数は、一千七百五十億。
GPT-2から桁違いに大きくなった。
訓練には、Common Crawlをはじめとするウェブ文書、書籍、Wikipediaなど、大量のテキストが用いられた。
投入された計算資源も、当時としては前例の少ない規模だった。
同じ時期、OpenAIではスケーリング則の研究が進んでいた。
大きくすれば、どのように良くなるのか。
どこまで予測できるのか。
GPT-3の建造と、スケーリングの海図は、互いに影響しながら進んでいた。
そして——研究者たちは、新しい岸辺を見た。
GPT-3は、一つのモデルで、文章を書く。翻訳する。要約する。質問に答える。コードの例を生成する。短い詩を作る。
しかも、多くの課題で、課題ごとの追加学習を行わなくても、指示や少数の例を文脈に置くだけで振る舞いを変えた。
論文はそれを、Few-shot Learning——少例学習と呼んだ。
実際には、モデルの内部重みがその場で書き換わるわけではない。
訓練が新たに始まるわけでもない。
与えられた文脈の中から、課題の形を読み取り、その続きとして答えるのである。
だが、外から見ると、それは例題を見て課題の形式を掴むように見えた。
この見え方が、研究者たちを驚かせた。
従来、多くのAIシステムは、課題ごとに専用の訓練を必要とした。
GPT-3は、少なくとも一部の課題で、その境界を曖昧にした。
一つの大きなモデルが、文脈に応じて多くの仕事をこなす。
それは、言語モデルという名の下に隠れていた、新しい汎用性の発見であった。
もちろん、そこは楽園ではなかった。
GPT-3は誤る。
存在しない事実を語る。
自信ありげに虚構を紡ぐ。
長い推論では道に迷う。
それでも、研究者たちは、新大陸の浜辺に立っていた。
砂はまだ粗く、森の奥は暗かった。
だが、そこがただの島ではないことだけは、すでに見え始めていた。

終 まだ、誰も知らなかった

コロンブスが新大陸に上陸した時、彼は、自分がインドに着いたと信じていた。
それが、まったく新しい大陸——後に『アメリカ』と名付けられる、新世界——であることを、彼は最後まで認めなかった。GPT-3もまた、似た運命を持っていた。OpenAIの研究者たちは、自分たちの発見を慎重に表現した。
『言語モデルは少例学習者である』。
それは正確な題名であった。
だが、控えめでもあった。
彼らが見つけたものは、より良い言語モデルだけではなかった。
人間が自然言語で機械に仕事を頼む、という新しい関係の入口であった。
二〇二〇年六月、GPT-3は開発者向けの限定的なAPIとして提供され始めた。モデルの重みそのものは公開されず、利用者はOpenAIの門を通じて、その力を試すことになった。
研究者、技術者、起業家たちが、次々と触れたものの、まだ大衆のものではなかった。
高価で、限定されたAPIであり、誰もが日常的に使う製品ではなかった。新大陸の存在は船乗りたちの間で語られていたが、世界の大半はまだ、その海岸線を見ていなかった。OpenAIの中では、議論が、続いていた。
『これを、もっと多くの人に、使ってもらうべきだろうか』
『いや、もう少し、慎重に進めるべきだ』
『だが、世間の関心が高まる前に、我々から発表しないと——』
これらの議論の先に——二〇二二年十一月三十日——一つの製品が、世に出ることになる。ChatGPT。
それは、誰にでも使える、対話型の言語モデル。
二〇二〇年の時点で、OpenAIの研究員たちは、まだ、自分たちが新大陸に上陸したばかりであった。
岸辺で、彼らは、立っていた。
向こうに、何が広がっているのか——まだ、誰も知らなかった。
第九章 了

第十章 プロメテウスの火

ChatGPTの衝撃(二〇二二)

一 二〇二二年・炉の手前

二〇二二年、サンフランシスコ。
すでにGPT-3は世に出ていた。
だが、それはまだ、誰もが自然に使える対話相手ではなかった。
APIを通じてモデルを使うことはできた。開発者たちは、その力を試していた。
それでも、大多数の人々にとって、大規模言語モデルは研究所や企業の奥にある、不思議な装置だった。
その距離を縮めるために必要だったのは、単にモデルを大きくすることではなかった。
人間の指示に、もっと素直に従うこと。
人間が望む答え方に、近づくこと。
その橋として、一つの重要な仕事があった。
InstructGPT。
二〇二二年一月、OpenAIは、人間のフィードバックを用いて、GPT-3を人間の指示に従いやすく調整した研究を公表した。
ロング・オウヤンらのチームが示したのは、巨大なモデルにさらに知識を詰め込むことだけが進歩ではない、ということだった。
同じ力でも、どう引き出すかで、使い勝手は変わる。
問いに答える。
指示に従う。
有害な出力を減らす。
人間の意図に、より近づく。
InstructGPTは、GPT-3からChatGPTへ向かう橋だった。
OpenAI自身も、後にChatGPTを、InstructGPTの『兄弟モデル』と説明することになる。
火は、すでにあった。
必要だったのは、その火を、人間が近づける炉に入れることだった。

二 RLHFという技

RLHF。
Reinforcement Learning from Human Feedback。
人間のフィードバックによる強化学習。
ChatGPTへ続く道の中心にあった技術である。
問題は明らかだった。
GPT-3は、膨大なテキストを読み、次に来る語を予測するよう訓練されていた。
その能力は驚異的であった。
だが、それは『人間の指示に従う』ことそのものを目的にした訓練ではない。
文章の続きを作る力と、役に立つ助手として答える力は、同じではなかった。
そこでOpenAIは、InstructGPTで用いた方法を、対話の形へ発展させていく。
おおまかに言えば、流れは三つである。
第一段階。
人間の訓練者が、望ましい応答の例を作り、それを使ってモデルを教師あり学習で微調整する。
第二段階。
一つの問いに対して複数のモデル回答を作り、人間がそれらを順位づけする。その比較データから、人間の好みを予測する『報酬モデル』を学習させる。
第三段階。
その報酬モデルを手がかりに、強化学習によって言語モデルをさらに調整する。
OpenAIは、この段階でPPOと呼ばれる手法を用いた。
ここで重要なのは、RLHFが機械に『真実』そのものを注入する魔法ではない、ということである。
人間がより望ましいと評価した振る舞いへ、モデルを近づける。
そのための仕組みである。
それでも、変化は大きかった。
問いに、問いとして答えやすくなる。
指示に従いやすくなる。
不適切な要求を拒むよう調整できる。
会話の文脈に合わせやすくなる。
本書の比喩でいえば、これは巨大な言語モデルの『社会化』に近い。
ただし、人間の社会化が完全でないように、RLHFも完全ではない。
人間の評価の偏りを受ける。
もっともらしい誤りは残る。
安全性と有用性のあいだで、難しい調整が必要になる。
それでも、野生の火は、炉の火へ近づいた。
二〇二二年初めまでに訓練されたGPT-3.5系列のモデルをもとに、対話用の微調整が進められた。
そして、機械は、会話を始める準備を整えていった。

三 十一月三十日

二〇二二年十一月三十日。
OpenAIは、ひとつの対話型AIを公開した。
名前は、ChatGPT。
ChatとGPTを組み合わせただけの、驚くほど飾り気のない名前だった。
発表もまた、巨大な式典ではなかった。
研究プレビューとして公開し、利用者からのフィードバックを集める。
OpenAIの告知は、その程度の慎重なものだった。
だが、公開されたものは、GPT-3の単なる入力欄ではなかった。
前の会話を受けて答える。
追加の質問に応じる。
誤りを指摘されれば修正を試みる。
不適切な要求には拒否を返す。
それまで研究論文やAPIの向こうにあった大規模言語モデルが、会話という最も古い人間の形式で、目の前に現れた。
そこから先、世界は、開発した側の想像を超える速度で反応することになる。
紀元前のギリシャ神話。
神々は火を独占していた。人間は暗く、寒く、弱かった。
プロメテウスは、その火を盗み、人類へ渡した。
そのとき、プロメテウス自身も、神々も、人類も——その火が、どれほどの力を持つかを、まだ十分には理解していなかった。
火は、人類が料理を作ることを可能にし、暖を取ることを可能にし、金属を加工することを可能にし、文明を築くことを可能にした。
プロメテウスが渡したのは、ただの『火』ではなかった。
それは、人類の在り方そのものを変える力であった。
二〇二二年十一月三十日。
人々の手元に渡ったのも、単なる新製品ではなかった。
言葉を通じて、誰もが直接触れられるAIだった。
火は、小さな研究プレビューとして放たれた。
そして、燃え広がった。

四 五日で百万人

ChatGPTを、最初に触ったのは、技術者たちであった。
彼らは、これを試した。質問してみる。詩を書かせる。コードを書かせる。文章を要約させる。
そして——驚いた。
『これは、本当にAIなのか』 『人間と話しているような感覚だ』 『これまでのチャットボットとは、根本的に違う』
技術者たちは、SNSに、自分たちの体験を、投稿し始めた。
衝撃の声が、ネット上に、広がった。
口コミは、加速度的に、拡散した。
——五日後。ChatGPTの利用者数は、一〇〇万人を、突破した。
これは、消費者向け技術の歴史で、前代未聞の数字であった。Facebookが、その規模に達するのに、十か月かかった。 Instagramは、二か月半。 Netflixに至っては、三年半を要した。ChatGPTは——わずか、五日。OpenAIの内部は混乱した。
サーバーは悲鳴を上げた。負荷は予想を超えた。
研究プレビューは、研究者たちの手を離れ、大衆のものになり始めていた。
アルトマン自身、公開から五日目に百万ユーザーの突破を短く投稿し、その数日後には、こうも書いている。
『ChatGPTはきわめて限定的だ。だが、いくつかのことが十分に得意なせいで、偉大さについての誤った印象を与えてしまう』
醒めた言葉であった。だが実際には、彼らは、自分たちが何を解き放ったのか、まだ完全には理解していなかった。
加速は、止まらなかった。
二〇二三年一月、公開からわずか二か月で、ChatGPTの月間利用者は一億人に達したと推計された。
よく引かれる比較では、電話が一億人規模に届くまでに、七十五年。
携帯電話は十六年。
インターネットは七年。Instagramは二年半、TikTokでも九か月。ChatGPTは、二か月。
これは、その時点で人類が手にした、最も急速に普及した技術であった。
学生、教師、医師、弁護士、芸術家、技術者、主婦、引退した老人——あらゆる年齢、あらゆる職業の人々が、自分の手元の機械で、AIと、文字通り対話を始めた。
プロメテウスの火は、ひそかに、世界中に、燃え広がっていた。
そして、これは、止まらない火、であった。

五 哲学者の問い

ChatGPTの登場は、技術的事件であるだけではなかった。
それは、文明的事件であった。
世界中の知識人が、これに、反応し始めた。
哲学者たちは、問うた——
『これは、知能なのか』
機械が、人間と区別がつかないほど自然な対話をし、人間が思いつかぬほど創造的な文章を書き、人間より丁寧で精確な要約を行う——これを、『知能』と呼ばずに、何と呼ぶのか。
七十年以上前、アラン・チューリングが、雑誌『マインド』に投稿した、あの論文——『計算機械と知能』。彼が予言した『考える機械』が、ここにいた。
教育者たちは、問うた——
『子供たちに、何を教えるべきか』
学生たちは、ChatGPTに、宿題をやらせるようになった。エッセイを書かせる、問題を解かせる、論文を書かせる。これを、どう扱うべきか。
『AIを使わせない』と禁じるのか。それとも、『AIを上手に使う』ことを、新しい教養として教えるのか。
各国の大学、各国の中等学校、各国の小学校——教育の現場で、議論が、激化した。
企業の経営者たちは、問うた——
『我々の仕事は、どう変わるのか』
文章を書く仕事、コードを書く仕事、コールセンターでの応対、翻訳、要約、相談——これらのすべてが、AIによって、ある程度、こなせるようになった。
雇用は、どう変わるのか。給与は。職業の未来は。
法律家たちは、問うた——
『著作権は、どうなるのか』ChatGPTは、インターネット中のテキストを、何十億単語と読んで、学習した。その中には、著作権で保護された作品が、無数に含まれていた。AIが学習に使うこと、AIが新しい文章を生成すること——これら一つ一つの法的位置は、どうなるのか。
世界中の出版社、新聞社、作家団体が、訴訟を起こし始めた。
そして——多くの普通の人々は、もっと素朴な問いを、抱いた——
『これは、本当に、私を理解しているのか』ChatGPTに、自分の悩みを打ち明けた人。子供が、ChatGPTに、宿題の相談をした。引退した老人が、ChatGPTと、寂しさを紛らわすために会話を続けた——
彼らは、感じていた。
『これは、本当に、機械なのか。それとも、何か、別のものなのか』
これらすべての問いに、まだ、答えは出ていない。

六 三国志の幕開け

ChatGPTの登場は、業界の勢力図を一変させた。
長らくAI研究の盟主であったGoogleは、内部で『コード・レッド(緊急事態)』を発令した。スンダル・ピチャイCEOは、半ば引退状態であったセルゲイ・ブリン共同創立者を、職場に呼び戻した。Googleが、こんなに慌てたことは、過去ほとんどなかった。Googleにとって、これは屈辱に近い衝撃であった。
トランスフォーマーを生んだのはGoogleである。BERTを生んだのもGoogleである。TPUを築き、世界最高峰の研究者を抱えていたのもGoogleである。
だが、世界の人々に最初に火を渡したのは、OpenAIだった。
数か月後、Googleは、対抗AI『Bard』を、急遽世に出した。だが、その性能はChatGPTに及ばず、業界では失笑を呼んだ。
そして、OpenAIの内部では——亀裂が、決定的になっていた。
ダリオ・アモデイらの『安全派』は、二〇二〇年末にOpenAIを離れ、翌二一年、別の結社『Anthropic』を立ち上げていた。ChatGPTの公開を、彼らはサンフランシスコの別の場所から、見つめていた。
強力な火は、ただ配ればよいものではない。
炉を設け、柵を作り、制御の技を磨かねばならない。
二〇二三年三月、AnthropicはClaudeを公開した。
『憲法AI』を謳い、慎重さ、安全性、対話の品位を前面に出したそのモデルは、ChatGPTの最大の競合の一つとなっていく。Metaは、別の道を選んだ。LLaMAをはじめとするモデルを公開し、開放路線へ踏み出した。巨大企業がすべてを閉じ込めるのではなく、研究者や開発者の手に重みを渡す。火を独占するのではなく、火種をばらまく戦略であった。
東方では、中国の研究勢力が追い上げた。
やがてDeepSeekのような新しい結社が、限られた資源で驚くべき効率を示し、米国の巨大モデルに迫っていく。
そして、イーロン・マスクはxAIを立ち上げた。
かつてOpenAIの創設に関わった人物が、別の旗を掲げて同じ戦場へ戻ってきたのである。
火は、ひとつではなくなった。OpenAI、Google、Anthropic、Meta、Microsoft、中国勢、xAI。
それぞれが別の炉を築き、別の火を掲げ、別の未来を語り始めた。
ここから、現代AI動乱史が始まった。
だが、その発火点は明らかであった。
二〇二二年十一月三十日。ChatGPTの公開。
あの日、小さく放たれた火が、すべての陣営を動かしたのである。

終 八十六年の弧

ここで、物語の始まりへ戻ろう。
一九三六年、ケンブリッジ。
二十四歳を目前にしたアラン・チューリングは、紙テープの上を動く抽象機械を描いた。
それは、まだ現実の機械ではなかった。
計算とは何かを問うための、ひとつの美しい思考実験であった。
それから——八十六年。
二〇二二年十一月三十日。
世界中の人々が、自分の手の中の小さな機械で、言葉を返すAIと対話する時代が来た。
一九五〇年、チューリングが『機械は考えることができるか』という問いを置き直してからは、七十二年が経っていた。
彼は、二十世紀の終わりごろには、人々が『機械が考える』と語ることへの違和感は大きく薄れているだろうと見通した。
時期も形も、そのまま予言どおりではない。
だが、問いの向きは驚くほど未来を指していた。
その七十二年の間に、私たちは、神経網の冬を経て、AlphaGoの神の一手を見て、トランスフォーマーの来航を経て、GPT-3の新大陸を経て、InstructGPTという橋を渡り——ここに、たどり着いた。
物語に登場した者たちを、もう一度、思い出してみる。
チューリングは、半分の林檎を残して去った。
ローゼンブラットは、蝋の翼で空を目指し、海に落ちた。
冬の研究者たちは、隠れた灯火を守り続けた。
ヒントンの工房から、弟子たちが世界へ散った。
李世乭は、敗北の中で一勝を返した。
八人の使徒は、世界の各地へ散った。
カプランたちは、巨大なモデルの海に数式の海図を描いた。
そして、人間の評価を学ぶ技術が、野生の火を炉へ近づけた。
——彼らすべてが、それぞれの場所で、それぞれの闘いをし、それぞれの灯火を、次の世代に渡してきた。
その総和の一つが、二〇二二年十一月三十日に結実した。
紙テープ。
ダートマスの船。
蝋の翼。
冬の灯火。
黒船。
工房。
神の一手。
八人の使徒。
西への航海。
そして、炉に入れられた火。
それらすべての糸が編み込まれ、一つの炎となった。
ChatGPTは、誰か一人の発明ではない。
それは八十六年にわたる、多くの研究者、技術者、評価者、利用者の営みの上に現れた。
失敗し、誤解され、冬を耐え、敗北し、それでも灯火を渡し続けた者たちの総和であった。
そして、物語はここで終わらない。
二〇二二年以降も、GPT-4、Claude、Geminiなど、次の世代のモデルが現れ、AIは新しい能力と新しい問題を同時に見せ続けている。
私たちは、まだ、新大陸の岸辺に立っているにすぎない。
向こうに何が広がっているのか——まだ、誰も、はっきりとは見えていない。
幻覚の問題は、まだ消えていない。
教育のあり方は、模索されている。
著作権をめぐる争いも続いている。
雇用への影響も、まだ途中にある。
そして、最も大きな問い——
『人類は、機械と、どう共に生きていくのか』
これに、まだ、誰も、最終的な答えを持っていない。
歴史は、まさに今この瞬間も、私たちの目の前で書かれ続けているのである。
プロメテウスの火は、ひとたび人間の手に渡れば、もう、神々の手には戻らない。
その火を、私たちが、いかに使うか——
それは、この火を受け取った私たち全員が、これから書いていく物語である。
第十章 了

あとがき

火を受け取った私たちへ
ここまで読み進めてくださった読者に、まず感謝を申し上げたい。
ここで描こうとしたのは、人工知能の技術史である。だが、単なる論文名と年号の列挙にはしたくなかった。なぜなら、AIの歴史は、冷たい機械の歴史である前に、きわめて人間的な歴史だからである。
そこには、若きチューリングの孤独があった。
ダートマスに集った研究者たちの無謀な楽観があった。
ローゼンブラットの蝋の翼があり、AIの冬の中で灯火を守った蘭学者たちがいた。
黒船のように押し寄せた統計革命があり、カスパロフの敗北があった。
ヒントン工房の長い忍耐があり、李世乭の第七十八手があった。
八人の使徒が書いた一篇の論文があり、GPT-3という新大陸への航海があった。
そして最後に、ChatGPTというプロメテウスの火が、人類の手元に渡された。
振り返れば、この物語を貫いていたのは、勝利だけではない。
むしろ、敗北であった。
失敗した研究。
時代に早すぎた発明。
嘲笑された仮説。
資金を断たれた研究室。
海に消えた研究者。
人間が機械に敗れた盤上の沈黙。
けれど、その敗北の灰の下には、いつも次の火種が残っていた。
パーセプトロンの失速は、深層学習の伏線となった。
エキスパートシステムの限界は、統計的手法を呼び込んだ。Deep Blueへの敗北は、人間と機械の新しい協働を生んだ。AlphaGoの衝撃は、人間の創造性をもう一度問い直させた。GPT-2をめぐる不安は、AIの安全性という新しい時代の課題を浮かび上がらせた。
敗北は、終わりではなかった。
敗北は、次の創造の種であった。
全体で繰り返し使ってきた比喩——林檎、船、蝋の翼、蘭学、黒船、工房、神の一手、使徒、新大陸、そして火——は、史実を飾るためだけのものではない。
比喩は、地図である。
もちろん、地図は土地そのものではない。コロンブスの航海とGPT-3の開発は同じ出来事ではない。プロメテウスの神話とChatGPTの公開も、文字通り重なるものではない。
だが、地図がなければ、私たちは広大な土地の全体像を見失う。AIの歴史はあまりに速く、あまりに複雑である。論文、企業、国家、モデル、資金、倫理、訴訟、教育、雇用——それらが一気に押し寄せる。その渦の中で、比喩は、私たちが現在地を確かめるための海図となる。
この物語を書きながら、私は何度も同じことを考えた。
人工知能とは、いったい誰のものなのか。
研究者のものか。
企業のものか。
国家のものか。
投資家のものか。
それとも、使うすべての人々のものなのか。ChatGPT以後、その問いはもはや専門家だけのものではなくなった。学生も、教師も、医師も、弁護士も、作家も、技術者も、親も、子どもも、退職した老人も、みな同じ火の前に立っている。
この火を、どう使うのか。
恐れすぎれば、私たちは可能性を失う。
崇めすぎれば、私たちは判断力を失う。
軽んじれば、火傷を負う。
独占すれば、争いが起こる。
野放しにすれば、森が燃える。
必要なのは、恐怖でも信仰でもなく、知恵である。AIは、人間の代わりになるためだけに現れたのではない。
人間を試すために現れたのだと思う。
私たちは、何を学ぶのか。
何を創るのか。
何を人間の仕事として残すのか。
何を機械に任せるのか。
そして、機械が何かをできるようになった後で、人間は何を望むのか。
最終章でたどり着いた問いは、そこにある。
『人類は、機械とどう共に生きていくのか』
この問いに、まだ答えはない。
おそらく、ひとつの正解もない。
だが、歴史を振り返ることには意味がある。
なぜなら、未来は突然生まれるのではなく、過去の無数の選択の上に立ち上がるからである。
チューリングの紙テープから、ChatGPTの生成まで。
その八十六年の弧は、ひとつの技術が完成へ向かった道ではない。人間が、自分自身の知性を外へ映し出そうとしてきた長い試みの軌跡であった。
そしてその試みは、まだ終わっていない。
むしろ、始まったばかりである。
プロメテウスの火は、もう神々の手には戻らない。
火は、私たちの手の中にある。
それで何を照らすのか。
何を鍛えるのか。
何を守り、何を燃やしてしまうのか。
その先の物語は、研究所の中だけで書かれるものではない。
企業の会議室だけで決まるものでもない。
国家の政策文書だけに閉じ込められるものでもない。
この火を使う、すべての人が書いていく物語である。
振り返れば、この物語を書きながら、私自身もまた、何かを学んでいた。
本書が、その長い物語を考えるための、ささやかな灯火になれば幸いである。

用語解説集

本書には、AI技術に関する多くの専門用語が登場します。文中では文学的表現を優先したため、用語の正確な定義は省略している箇所があります。本巻末解説集は、より深く本書をお読みになりたい方のための、簡潔な参考集です。

I. AIの基本概念

人工知能(Artificial Intelligence, AI) 人間の知的な活動——認識、理解、判断、創造、対話など——に関わる能力を機械で実現しようとする研究・技術の総称。「Artificial Intelligence」という名称は、一九五五年のダートマス研究計画提案書に記され、一九五六年の研究集会を通じて分野の旗印となった。本書第二章で詳述。
機械学習(Machine Learning) 明示的なルールを書くのではなく、データから機械自身がパターンを学ぶ手法。AI研究の主流的アプローチ。
ディープラーニング(Deep Learning, 深層学習) 多層のニューラルネットワークを用いる機械学習の総称。二〇〇六年のヒントンらの深層信念網研究は、深いネットワークへの関心を復興させる重要な契機となったが、「ディープラーニング」という語をヒントンがこの年に命名したわけではない。二〇一二年のAlexNet以降、AI研究の中心的潮流となった。
ニューラルネットワーク(Neural Network) 人間の脳の神経細胞の構造を模した数学的モデル。複数の「ノード(疑似神経細胞)」が層をなして接続される。本書第三章のローゼンブラットの「パーセプトロン」が、その最初期の実装の一つ。
パラメータ ニューラルネットワーク内部の調整可能な数値。学習を通じてこの数値が最適化され、機械の知能が形成される。GPT-3は一千七百五十億個のパラメータを持つ。

II. 機械学習の主要手法

記号主義AI(Symbolic AI) 人間の知識を論理規則として明示的に書き出し、機械に与える方式。一九六〇〜八〇年代の主流。エキスパートシステムが代表例。本書第四章「朱子学の春」で詳述。
統計的機械学習(Statistical Machine Learning) データから統計的な規則性や予測モデルを学ぶ手法群。一九九〇年代に大きく台頭し、記号主義と並ぶ、あるいは多くの領域で中心的なアプローチとなった。本書第五章で「黒船」として描かれる。
SVM(サポートベクターマシン, Support Vector Machine) ヴラジーミル・ヴァプニクが体系化した、データを分類する境界線を数学的に最適化する手法。一九九〇年代の代表的アルゴリズム。
ベイジアン・ネットワーク(Bayesian Network) ジューディア・パールが体系化した、確率的な因果関係を表現する手法。不確実性のある推論に有用。
HMM(隠れマルコフモデル, Hidden Markov Model) 時系列データを確率的に扱う手法。音声認識などで一九九〇年代に広く使われた。フレデリック・イェリネックらが応用を主導。
強化学習(Reinforcement Learning, RL) 機械が試行錯誤を通じて、「報酬」を最大化する行動を学ぶ手法。AlphaGoの基盤技術の一つ。第七章で詳述。

III. ニューラルネットワーク関連

パーセプトロン(Perceptron) フランク・ローゼンブラットが一九五七年に考案した、最も単純なニューラルネットワーク。第三章の主人公。
XOR問題 単層の線形パーセプトロンでは「排他的論理和(AかBの一方だけが真) 」を表現できないという代表的な限界。『Perceptrons』(一九六九)はこうした限界を数学的に分析した。ニューラルネット研究の停滞は、この本だけでなく、計算資源、学習法、研究資金、過大な期待など複数の要因による。
バックプロパゲーション(Backpropagation, 誤差逆伝播法) 出力誤差を後段から前段へ伝え、各重みを調整することで多層ニューラルネットワークを訓練する方法。先行研究は存在するが、一九八六年のラメルハート、ヒントン、ウィリアムズの論文が、その有効性を広く知らしめた。第四章のクライマックスを担う。
CNN(畳み込みニューラルネットワーク, Convolutional Neural Network) 画像認識に特化したニューラルネットワーク構造。ヤン・ルカンが一九八〇年代から発展させた。AlexNet(二〇一二年)はこの構造を採用。
RNN(リカレントニューラルネット, Recurrent Neural Network) 時系列データ(文章、音声など)を順番に処理するニューラルネットワーク。Transformer登場まで自然言語処理の主流だった。
LSTM(Long Short-Term Memory) RNNの改良版。一九九〇年代に発明され、長い文脈を扱えるようにした。
深層信念網(Deep Belief Network) ヒントンが二〇〇六年に発表した、深い層を持つネットワーク構造。「ディープラーニング」復活のきっかけ。

IV. Transformer時代

Transformer(トランスフォーマー) 二〇一七年六月、八人の研究者が『Attention Is All You Need』で発表したニューラルネットワーク構造。再帰や畳み込みに頼らず、Attentionを中心に系列を処理する。現在の多くの大規模言語モデルや生成AIの基盤技術の一つ。
Attention(注意機構) ある要素を処理するとき、入力中のどの要素をどれだけ強く参照するかを重みとして計算する仕組み。Transformer以前から用いられ、TransformerではSelf-Attentionが中核となった。
Self-Attention(自己注意) 文章内のすべての単語が、互いに注意を向け合う仕組み。並列計算が可能で、遠い単語同士でも直接関係づけられる。
事前学習(Pre-training) 大量のテキストで、機械に言語の一般的な構造を学ばせる第一段階の学習。
微調整(Fine-tuning) 事前学習済みモデルを、特定の課題や望ましい振る舞いに合わせて追加学習すること。GPT-1は「大規模な事前学習の後、下流課題に微調整する」という流れを示した代表的な初期研究の一つ。
スケーリング則(Scaling Laws) ジャレッド・カプランらが二〇二〇年に示した、言語モデルの損失とモデル規模・データ量・計算量のあいだに広い範囲でべき乗則的な関係が見られるという経験則。大規模化の効果を予測するための重要な手がかりとなった。
Few-shot Learning(少例学習) わずか数個の例だけを示して、新しい課題を機械にこなさせる手法。GPT-3が驚くべき能力を見せた領域。
大規模言語モデル(Large Language Model, LLM) 莫大なパラメータ数を持つTransformerベースの言語モデル。GPT、Claude、Geminiなどが代表例。
RLHF(人間のフィードバックによる強化学習, Reinforcement Learning from Human Feedback) 人間の評価者がAIの出力をランク付けし、機械が「人間が好む答え方」を学ぶ訓練法。ChatGPTの中核技術。
幻覚(ハルシネーション, Hallucination) AIが「もっともらしいが事実ではない」内容を、自信を持って出力する現象。現代AIの最大の課題の一つ。

V. 計算インフラ

GPU(Graphics Processing Unit, グラフィックス処理装置) もとはゲーマー向けの画像処理用チップ。並列計算に優れ、ニューラルネットワークの訓練に転用されて、現代AIの計算基盤となった。エヌビディア社が業界の盟主。
CUDA エヌビディア社が二〇〇七年に公開した、GPUを汎用計算に使うための仕組み。
TPU(Tensor Processing Unit) Googleが独自開発した、AI計算専用のチップ。GPU の代替・並列利用される。

VI. 生成AI関連

生成AI(Generative AI) 新しいテキスト、画像、音楽、動画などを「生成」できるAI。ChatGPTの登場以降、急速に普及した。
GAN(敵対的生成ネットワーク, Generative Adversarial Network) 二〇一四年にイアン・グッドフェローらが発表した生成モデル。生成器と識別器を競わせることで学習し、深層学習による画像生成を大きく前進させた。
拡散モデル(Diffusion Model) データへ段階的にノイズを加える過程と、その逆向きのノイズ除去過程を学ぶ生成モデル。Stable Diffusionは、潜在空間で拡散過程を扱うLatent Diffusion Modelを基盤とする。
マルチモーダルAI(Multimodal AI) テキスト・画像・音声・動画など、複数の形式の情報を同時に扱えるAI。GPT-4、Claude 3などが該当。

VII. AI研究の重要組織

OpenAI 二〇一五年、サム・アルトマン、イーロン・マスク、イリヤ・サツケヴァーらが創立。当初は非営利、後に営利化。GPTシリーズとChatGPTで知られる。
Anthropic 二〇二一年、ダリオ・アモデイ、ダニエラ・アモデイらが、OpenAIから離脱して創立。AIの安全性を最優先する立場。Claudeを開発。第十章で詳述。
DeepMind 二〇一〇年、デミス・ハサビスらがロンドンで創立。二〇一四年にGoogleに買収。AlphaGo、AlphaFoldで知られる。
Google Brain 二〇一一年、アンドリュー・ンとジェフ・ディーンがGoogle内に設立。Transformer論文の故郷。
Meta AI(旧Facebook AI Research / FAIR) Meta社のAI研究部門。ヤン・ルカンが長く主任科学者を務めた(二〇二五年末に退社)。Llamaシリーズの開発で知られる。
xAI 二〇二三年、イーロン・マスクが創立。「Grok」を開発。
DeepSeek 中国の幻方量化(ヘッジファンド)が出資する、北京・杭州拠点のAI研究機関。極めて効率的なモデル開発で知られる。
Sakana AI 二〇二三年、リオン・ジョーンズ(Transformer論文共著者)、David Ha、伊藤錬が東京で創立。自然界に学ぶAIを志向。
ICOT(新世代コンピュータ技術開発機構) 日本の第五世代コンピュータ計画(一九八二-一九九二)の中核機関。所長は渕一博博士。
DARPA / ARPA(米国防高等研究計画局) 米国国防総省の研究機関。AI研究への初期資金の多くを提供した。

VIII. 主要モデル・論文(時系列)

ダートマス会議の提案書(一九五五) ジョン・マッカーシー、マービン・ミンスキー、ナサニエル・ロチェスター、クロード・シャノンがまとめた研究計画書。「Artificial Intelligence」という名称が分野の旗印として用いられた初期の重要文書。
論理理論機(Logic Theorist, 一九五六) ニューウェル、サイモン、ショーによる、最初期のAIプログラムの代表。プリンキピア・マテマティカの定理を証明できた。
パーセプトロン(Mark I, 一九六〇) フランク・ローゼンブラットらがコーネル航空研究所で構築した、最初期の本格的ニューラルネットワーク・マシン。構想は一九五七年に発表され、実機は一九六〇年に公開された。
『Perceptrons』(一九六九) マービン・ミンスキーとシーモア・パパートによる、パーセプトロンの能力と限界を数学的に分析した本。後世、ニューラルネット研究停滞の象徴として語られるが、停滞の原因は一冊の本だけではない。
DENDRAL(一九六〇〜七〇年代) スタンフォード大学のエドワード・ファイゲンバウムらによる、最初期のエキスパートシステム。化学構造の推定を行う。
MYCIN(一九七〇年代後半) エドワード・ショートリフによる、感染症診断のエキスパートシステム。約五百個のルールを持ち、専門医に匹敵する精度を達成。
バックプロパゲーション論文(一九八六) デイヴィッド・ラメルハート、ジェフリー・ヒントン、ロナルド・ウィリアムズによる『Learning representations by back-propagating errors』。先行研究のあった誤差逆伝播法を、多層ニューラルネットの有力な学習法として広く普及させた。
Deep Blue(一九九七) IBMが開発したチェス専用機。世界チャンピオン、ガリー・カスパロフを破った。
深層信念網論文(二〇〇六) ジェフリー・ヒントン、サイモン・オシンデロ、イー・ワイ・テーによる『A fast learning algorithm for deep belief nets』。深いニューラルネットワーク研究の復興を象徴する重要論文。
ImageNet(二〇〇九) フェイフェイ・リーが構築した、大規模画像データベース。公開時は約三百二十万枚、のちに一千四百万枚超へ成長。
AlexNet(二〇一二) クリジェフスキー、サツケヴァー、ヒントンによる深層CNN。ImageNet Challengeで圧勝。
DQN(二〇一三) DeepMindによる深層強化学習の初期研究。生の画面ピクセルを入力として七つのAtari 2600ゲームを学習し、その後の研究で対象ゲームを拡大した。
AlphaGo(二〇一六) DeepMindによる囲碁AI。李世乭九段を破った。
『Attention Is All You Need』(二〇一七) アシシュ・ヴァスワニら八人によるTransformer論文。Attentionそのものには先行研究があるが、再帰や畳み込みを使わず、Attentionを中心に構成した点が画期的だった。
BERT(二〇一八) Google・ジェイコブ・デブリンらによる、双方向Transformer。
GPT-1(二〇一八)/GPT-2(二〇一九)/GPT-3(二〇二〇) OpenAIによる大規模言語モデル。アレック・ラドフォードらが筆頭著者。
スケーリング則論文(二〇二〇) ジャレッド・カプランらによる、規模と性能の関係の数式化。
ChatGPT(二〇二二年十一月三十日) OpenAIが研究プレビューとして公開した対話型AI。GPT-3.5系列のモデルをもとに、InstructGPTと同系統のRLHF手法を用いて対話向けに調整された。
Claude(二〇二三年三月) Anthropicによる対話型AI。安全性を重視。
Bard / Gemini(二〇二三-二〇二四) GoogleによるChatGPT対抗のAI。

IX. 主要な歴史的概念

チューリング・テスト(一九五〇) アラン・チューリングが提唱した、機械の知能を判定する思考実験。文章だけの会話で、人間と機械を見分けられないなら、その機械は「考えている」と認める。
ダートマス会議(一九五六) 米国ニューハンプシャー州ダートマス大学にて開催された、AI研究の出発点となった研究集会。
AIの冬(AI Winter) AIへの期待が後退し、研究資金や社会的関心が縮小した時期を指す通称。第一次・第二次と区分されることが多いが、開始・終了年や原因の捉え方には幅があり、複数の要因が重なって起きた。
ライトヒル報告書(一九七三) 英国の数学者ジェームズ・ライトヒルがまとめたAI研究評価。英国におけるAI研究支援の大幅縮小に影響し、第一次AI冬を象徴する出来事の一つとなった。
第五世代コンピュータ計画(一九八二-一九九二) 日本の通産省が主導した、AI国家プロジェクト。渕一博博士が中心。第四章で詳述。
AI安全性(AI Safety) 強力なAIが人類にもたらす潜在的危険を研究し、制御する分野。Anthropicの設立理念の核心。
アライメント(Alignment) AIの目標を、人類の価値観や意図と「整合」させる課題。AI安全性研究の中心問題。

X. 主要人物の補足

本書本文では文学的描写を優先したため、ここで簡潔な経歴を補足する。
アラン・チューリング(一九一二—一九五四) 英国の数学者。AIの理論的基礎を築いた。第一章。
ジョン・マッカーシー(一九二七—二〇一一) 米国の計算機学者。「人工知能」の命名者。第二章。
マービン・ミンスキー(一九二七—二〇一六) 米国のAI研究者。MIT人工知能研究所創設者の一人。
ハーバート・サイモン(一九一六—二〇〇一) 米国の認知科学者。一九七八年ノーベル経済学賞。
アレン・ニューウェル(一九二七—一九九二) 米国の計算機学者。サイモンの共同研究者。
クロード・シャノン(一九一六—二〇〇一) 米国の数学者。情報理論の創始者。
フランク・ローゼンブラット(一九二八—一九七一) 米国の心理学者。パーセプトロンの発明者。第三章。
ジェフリー・ヒントン(一九四七—) 英国系カナダ人。ディープラーニングの父。二〇一八年チューリング賞、二〇二四年ノーベル物理学賞。
ヤン・ルカン(一九六〇—) フランス系米国人。CNN発展の中心人物。元Meta AI主任科学者(二〇二五年末退社)。二〇一八年チューリング賞。
ヨシュア・ベンジオ(一九六四—) カナダ人。モントリオール大学教授。二〇一八年チューリング賞。
フェイフェイ・リー(一九七六—) 中国系米国人。ImageNet創設者。「AIの母」とも呼ばれる。
アレックス・クリジェフスキー ウクライナ系カナダ人。AlexNetの主実装者。
イリヤ・サツケヴァー(一九八六—) ロシア系イスラエル系カナダ人。OpenAIの主任研究員を経て、SSI創立。
渕一博(一九三六—二〇〇六) 日本の計算機学者。第五世代コンピュータ計画の中心人物。
ヴラジーミル・ヴァプニク(一九三六—) ロシア生まれの統計学者。SVMの理論的体系化者。
ジューディア・パール(一九三六—) イスラエル系米国人。ベイジアンネットワークの確立者。二〇一一年チューリング賞。
ガリー・カスパロフ(一九六三—) ソ連→ロシアのチェス選手。世界選手権者。Deep Blueに敗北。
デミス・ハサビス(一九七六—) 英国人。DeepMind創立者。二〇二四年ノーベル化学賞。
李世乭(イ・セドル)(一九八三—) 韓国の囲碁棋士。二〇一六年の五番勝負でAlphaGoから唯一の一勝を挙げた。二〇一九年引退。
サム・アルトマン(一九八五—) 米国の起業家。OpenAI CEO。
ダリオ・アモデイ(一九八三—) 米国の物理学者・AI研究者。Anthropic共同創立者・CEO。
イーロン・マスク(一九七一—) 南アフリカ系米国の起業家。OpenAI共同創立者、後に離脱してxAI創立。
リオン・ジョーンズ ウェールズ生まれ。Transformer論文共著者。Sakana AI共同創立者。

XI. 補足:AI研究の三つの大波

本書を通読する際の見取り図として、AI研究には大きく三つの「波」があったことを記しておく。
第一の波(一九五〇〜八〇年代) :記号主義AI。論理規則とエキスパートシステムの時代。第二章〜第四章で扱う。
第二の波(一九九〇〜二〇〇〇年代) :統計的機械学習。SVM、ベイズ、HMMなどの確率的手法の時代。第五章で扱う。
第三の波(二〇一二年〜現在) :ディープラーニング。多層ニューラルネットワークと大規模言語モデルの時代。第六章〜第十章で扱う。これら三つの波は、それぞれが前の波の限界を超えるために生まれた。そして第三の波は、二〇二二年のChatGPT登場により、人類全体の日常へと降臨した。

English version

Heirs to the Fire of Prometheus
From Turing to ChatGPT: An Eighty-Six-Year Story of AI


About This Book

This book is a work of narrative history, built on papers, public records, and historical sources, that traces the development of AI as a story of people. The gestures of its figures and the air of its scenes include literary description reconstructed from the historical record. Conversations and inner thoughts for which no verbatim record survives are treated not as verified statements, but as reconstructions, made to convey the meaning of a scene.

Prologue

From the earliest times, humanity has dreamed of making a likeness of itself.
From clay.
From mud.
From ivory.
And sometimes, from the flesh of the dead.
The myths and stories of the world tell of that dream.
The ivory statue carved by Pygmalion drew breath by the grace of the goddess Aphrodite. The Rabbi of Prague shaped a golem from clay and gave it life by carving the word “Truth” upon its brow. In a nineteenth-century tale, a young natural philosopher born in Switzerland stitched together flesh gathered from graveyards and made a living thing.
Why has humanity told such stories, again and again?
Perhaps, through the dream of making another self, we have gone on asking the mystery of what it is to “think,” and what it is to “live.”
What is life?
What is mind?
Can we make a thing that speaks, that answers, that chooses, that errs, that suffers doubt?
When the made thing opens its eyes, what is called into question is not only the true nature of the monster or the doll. We who made it are asked, in turn, what manner of beings we are.
In the twentieth century, that dream left the forge of myth and became a problem of mathematics and engineering.
1936, Cambridge.
A young mathematician named Alan Turing, not yet twenty-four, drew the concept of a calculating machine with nothing but paper, symbols, and a single tape. It had, as yet, no metal body, no voice, no face. But deep within that abstract machine lay the prototype of what would later be called the “thinking machine.”
From there — eighty-six years.
30 November 2022.
People all over the world began, quite literally, to converse with an AI that returned words to them, through the small machines in their hands.
This is the story of those eighty-six years.
It is not a story of victories alone.
It is a story of struggle, a story of defeat, a story of the resurrection that comes after a long winter.
One of them departed, leaving half an apple behind.
One challenged the sky on wings of wax, and fell into the sea.
One kept the lamp of a laboratory burning through the winter, refusing to let it go out.
One champion was defeated by a machine upon the board.
One player answered with a move that went beyond the machine's reading.
Eight of them left a single paper behind, and scattered across the world.
And then, one day, someone stole the fire of the gods and passed it, in secret, to humanity.
Not one of those who appear in this story understood completely, at the time, what it was they were making.
And yet they went on.
One step, and another.
Into the dark sea.
Into the unseen sky.
Toward a shore not yet mapped.
And now we stand upon that shore.
What lies beyond — no one, yet, knows.
This is that story.

Chapter One: Half an Apple

A Portrait of Alan Turing (1912–1954)

1. A Bicycle in Cambridge

May 1936, Cambridge.
Down a narrow lane running behind King's College, a young man was riding a bicycle. Tall and thin. Hair disordered by the wind. A fountain pen in his breast pocket. The fingers gripping the handlebars were long and slender, and looked somehow nervous.
Alan Mathison Turing. A young mathematician, not yet twenty-four. The year before, at only twenty-two, he had been elected a Fellow of King's College. Everyone acknowledged his talent; but where that talent was heading, no one could yet see.
The petals of late spring lay scattered on the lawns. He hardly saw them. Inside his head, a machine that did not yet exist in this world was running.
A long paper tape.
A single read-write head, crawling along it.
It reads a symbol, changes its state, writes a symbol, and moves one square to the right or left.
That, and nothing more.
And yet within that “nothing more,” he saw the skeleton of the act of computation. If you pared down, to the very limit, the calculation a human being performs with paper and pencil, what would remain? Reading symbols. Following rules. Passing to the next state. Then computation was not the lightning-flash of genius — it was something that could be written down as procedure.
He decided to call this abstract machine an “automatic machine.”
The title of the paper was already decided.
“On Computable Numbers, with an Application to the Entscheidungsproblem.”
To most people of the day, it was no more than an abstruse argument exchanged in some inner room of mathematics. But a later age would come to know: what was drawn there was the prototype of the modern computer — and, farther off still, the first outline of the dream called artificial intelligence.
As Galileo had looked through his telescope into the heavens and seen that the Earth was not the center of the universe, Turing looked into the hands of one who computes, and saw that a part of intelligence could be reduced to procedure.
The discovery did not yet threaten anyone.
It was not yet heresy.
The Turing of that moment could hardly have imagined that the small machine in his head would one day come to dwell on every desk in the world, in factories, in hospitals, in weapons, and in the palms of human hands.
He was simply looking at a beautiful picture.
One tape, and one head.
A quiet picture that, with only that, would rewrite the world.

2. The Huts at Bletchley

Three years later, the world was at war.
Buckinghamshire. Bletchley Park. Around the old manor house, makeshift wooden huts had been thrown up. Elegant lawns and damp plank walls. In that strange combination, mathematicians, linguists, chess enthusiasts, and champions of the crossword puzzle, gathered from all over Britain, fought with pencils, paper, and tea.
Their adversary was the German cipher.
Turing's post was Hut 8, the section that handled naval codes.
In the North Atlantic, the U-boats of Nazi Germany were sinking, one after another, the merchant ships bound for Britain. Food, fuel, weapons — nothing could reach Britain except across the sea. To command that sea, the German forces used a cipher machine called Enigma.
Enigma was a beautiful machine. Multiple rotors reconfigured its electric circuits, and the settings changed with every day. The possible combinations reached astronomical numbers. To try them all by human hand was, in practice, impossible.
But Turing stared at that “impossible” as a problem of logic.
Inheriting the breach opened by the Polish mathematicians, he and his colleagues conceived a machine to stand against Enigma — the device that would later be called the Bombe.
It was not a machine that thought of answers.
It was a machine that traced, without ever tiring, the gaps in logic that human beings had found.
Electric drums turned, testing possible settings at tremendous speed. A single day's trials were worth centuries of human labor. In time, the cipher that could not be broken began to break.
The Germans did not know. That their most secret communications were being read in huts in the English countryside. That orders exchanged beneath the waters of the Atlantic were being spread out, as slips of paper, on desks in wooden huts.
Postwar research would suggest that the codebreaking at Bletchley Park shortened the war by years and saved lives beyond counting. But that fact was long sealed away as a state secret.
The world did not know what Turing had done.
The world did not know, while he lived, that it had been saved by him.
His colleagues remembered his odd habits. How he chained his mug to the radiator so that it would not be stolen. How, his hay fever being severe, he wore a gas mask when he rode his bicycle. How, over long distances, he ran nearly fast enough to press an Olympic athlete.
A strange man, people said.
But at the same time, everyone knew.
Without that strange man, the war might have gone on far longer.

3. The Imitation Game

Five years passed after the war's end. 1950.
In the philosophical journal Mind, a single paper appeared. Its title was plain.
“Computing Machinery and Intelligence.”
At its opening, Turing asks:
“Can machines think?”
But he changed the shape of the question at once. Try to define what “thinking” is, and the argument runs on without end. What is mind? What is consciousness? What is the soul? Such questions were a labyrinth humanity had carried for more than a thousand years.
Turing did not enter that labyrinth by the front gate. Instead, he proposed a game.
In separate rooms, a human and a machine. A judge converses with both, through written words alone. No voice, no face. Only text passes back and forth. If the judge cannot tell the human from the machine — then may we not regard that machine as thinking?
A thought experiment later called the Turing Test.
His sharpness lay not in giving an answer. It lay in moving the place where the question stood. Not to worship intelligence as a mystery, but to treat it as observable behavior. Not to ask whether a soul resides within, but to make the responses that appear without the matter at issue.
It was an act that handed a question of philosophy over into the hands of engineering.
In the paper, he made a prediction: within some fifty years, people would be able to speak of “machines thinking” without expecting much resistance.
It took longer than that for the prediction to become the scenery of daily life. But the direction was not wrong. Some seventy years on, people would come to converse with machines, in writing, across small screens.
They would consult them, question them, joke with them, grow angry at them — even come to be encouraged by them.
Turing never saw that scene.

4. A Letter in February

The end of January 1952, Manchester.
One night, a young man stayed at Turing's house. His name was Arnold Murray. A few days later, the house was burgled. Turing reported it to the police.
Investigating the theft, the police arrived at another fact.
In the Britain of that day, relations between men were punished as “gross indecency.” Turing made no attempt to hide his relationship. To him, it was not a wrong that needed hiding.
The law did not see it that way.
Much of his wartime work still lay deep under state secrecy. The role he had played at Bletchley Park could not serve, in open court, as a medal to shield him. Service rendered to the nation and the law the nation had written were weighed on separate scales.
The trial came, and he was found guilty.
The sentence: prison, or the hormone injections then called “treatment.” He chose the latter. The administration of synthetic female hormones changed his body. His chest swelled; his strength declined. The body that had loved to run no longer moved as it once had.
Even so, the work went on.
In February, with the trial before him, in a letter to his friend Norman Routledge, he set down a syllogism.
Turing believes machines think.
Turing lies with men.
Therefore machines do not think.
It was a biting piece of wit.
A logician, he knew how coarse was the logic by which the world judged a man. And he took that coarseness, set it once more into the form of logic, and laughed it back.
Three centuries and more before, Galileo Galilei had been brought before the Inquisition and made to recant the heliocentric theory. He was judged not because he had moved the stars, but because he had spoken of what he saw exactly as he had seen it.
Turing was not judged for his thought about machines. But when authority sets its norms above reality, a man can be made a criminal merely for saying that he sees what he sees.
Legend has it that Galileo, leaving the court, murmured under his breath: “And yet it moves.”
What Turing murmured, the records do not tell.
Only the papers he wrote remain.
In them, machines had quietly begun to think.

5. Half an Apple

The night of 7 June 1954.
Wilmslow, Cheshire. Turing's house.
The next morning, his housekeeper found his body. He was forty-one.
Beside him lay an apple, half-eaten.
The examination found cyanide in his body. Officially, the death was ruled a suicide. Yet there is no record that the apple itself was ever tested for cyanide. He had been conducting chemical experiments at home, and might well have handled cyanide there. Voices would later be raised, pointing to the possibility of an accident.
Where does the truth lie?
Even so, the half apple remained.
Before it was ever evidence, it had taken on the aspect of something close to legend. Turing was fond of Disney's Snow White, and is said to have sung to himself the song from the scene in which the Queen brews the poisoned apple.
The world killed the man who had saved it.
Write it so, and the story is easy to understand.
But history is not made of easy stories alone. In his death, layer lies upon layer: the possibility of accident, the shadow of solitude, the violence of the institutions of his age.
Only one thing is certain.
The world did not pay him the respect he was owed.
His death was reported only briefly in the newspapers of the day. A mathematician had died at home. That was all. What he had achieved; how many lives had been saved by his work; how his abstract machine of 1936 would go on to change the world itself —
the world did not yet know.
Only the half apple remained, at his bedside.

6. Half a Century On

2009.
The British Prime Minister, Gordon Brown, issued an official apology to Turing on behalf of the government. He deserved, Brown said, so much better.
In 2013, Queen Elizabeth granted Turing a royal pardon.
In 2017, a retroactive pardon was extended to the tens of thousands of men who, like him, had been convicted for homosexual acts in the past — what came to be known as the “Alan Turing law.”
In 2021, his portrait was engraved on the British fifty-pound note.
The face of a man once branded a criminal became the currency of the nation.
For Galileo to receive what amounted to rehabilitation from the papacy took three hundred and fifty-nine years. Turing's honor was restored in little more than half a century.
Should we say: it took only half a century?
Or should we say: it took a full half century?
History does not right itself of its own accord. Someone remembers; someone asks again; someone demands an apology. At the end of that accumulation, the past turns, at last, a little.
Still — it turned.
The nation that had once judged him came, in time, to engrave his face upon its banknotes.
The world that had once hidden his name came, in time, to call him the father of computer science.

Coda: The Picture That Remained

What Turing left behind is not a machine.
It is a picture.
A long tape. Symbols. States. Rules. A head that reads, writes, and moves.
It was so simple as to be almost poor. And yet in that poverty dwelt everything of the modern computer. Computers, ciphers, programs, and artificial intelligence too — all are distant variations on this picture. ChatGPT, Claude, Gemini, AlphaGo — none of them runs on paper tape any longer. They run through circuits of silicon, cross vast seas of data, and choose their words and their moves out of a web of probabilities.
And still, at the root of it all, Turing's picture remains.
A thing that changes its state according to rules.
A thing that reads symbols, writes symbols, and moves on to the next step.
A thing that shines a light back, from the machine's side, upon the behavior human beings have called “thinking.”
He departed, leaving half an apple behind.
The half that remains is not only the riddle of his death.
What is human intelligence? How far can a machine think? How much of the thought we believed was ours alone will we come to share with machines?
Half of that question remains, even now.
Was it poison, or was it fruit?
Was it loss, or was it a beginning?
Turing's half apple sank to the bottom of the twentieth century, and rose again upon the palm of the twenty-first.
And still — do machines think?
The picture he left behind is moving quietly, even now.
End of Chapter One

Chapter Two: A Ship Without Oar or Rudder

The Summer of the Dartmouth Conference (1956)

1. The Proposal at Hanover

August 1955, Hanover.
A small town of a few thousand people, in a corner of New Hampshire. John McCarthy, a young assistant professor in the mathematics department of Dartmouth College, was about to turn twenty-eight.
At the beginning of that year, he had come to this country college on the East Coast from Stanford University, in California.
That summer, he set a name down on paper.
Artificial Intelligence.
An intelligence made by human hands.
At the head of the research proposal, drawn up under the date of 31 August 1955, stood four names: McCarthy, Marvin Minsky, Nathaniel Rochester, Claude Shannon. And there, written out plainly, was the term that had not yet taken root as a field of learning.
“Artificial intelligence.”
Similar words already existed. Cybernetics. Automatic control. Information theory. Complex information processing.
But what McCarthy wanted was a broader flag.
Not merely to make machines calculate.
To make machines perform intelligent behavior itself.
A name under which learning, language, abstraction, and problem-solving could all be gathered.
The proposal raised the plan of bringing some ten researchers together at Dartmouth for two months of the summer. At its center stood an astonishingly bold conjecture: that every aspect of intelligence could in principle be described so precisely that a machine could be made to simulate it.
Two months. Ten people. The unraveling of intelligence.
Looking back now, it is optimism verging on recklessness. But beginnings almost always run on optimism for fuel. Sometimes only those who have no chart can row out toward distant waters.
The plan did not proceed as first conceived. The grant approved by the Rockefeller Foundation came to five weeks' worth: seven thousand five hundred dollars. Even so, it was enough.
History often begins from small numbers such as these.
A single proposal.
A few weeks of summer.
Some ten researchers.
Seven thousand five hundred dollars.
And one name whose contents were not yet settled.
Artificial intelligence.

2. Arrivals in Summer

The summer of 1956, Dartmouth College.
The researchers gathered, little by little, in the town of Hanover. Not all of them assembled on the same day. It was not, in the first place, the kind of orderly international conference we imagine today — no finely fixed agenda, no allotted speaking times, no minutes put in order.
It was, rather, a gathering-place for a young field of learning.
Marvin Minsky. Twenty-eight. A young man possessed by the border between brain and machine.
Claude Shannon. Forty. The giant who had given the invisible thing called information a mathematical skeleton.
Nathaniel Rochester. A researcher at IBM, a practical man who had taken part in designing the most advanced computers of the day.
Arthur Samuel. Building a program that played checkers, he was probing the possibility that a machine might grow stronger from experience.
Ray Solomonoff, Oliver Selfridge, Trenchard More.
And Herbert Simon and Allen Newell.
The participants came and went at different times, and the discussion did not stay inside the room. It must have flowed out to the dining hall, onto the lawns, into the town at night.
What they spoke of was large.
Machine translation. Chess. Theorem proving. Learning. Neural circuits. Human problem-solving. And intelligence itself.
Around the AI researchers of that day hung an optimism: that within a few years, or a dozen and some, machines might catch up with much of human intellectual ability. Chess, translation, theorem proving — once recast as problems for the computer, they looked like mere matters of speed and ingenuity.
That estimate, of course, was naive.
But it would be cruel to demand of them, in that summer, that they foresee even the failures to come.
No one who begins an unknown science can measure its true difficulty from the start.
If they had known the real difficulty completely, perhaps no one would have rowed out at all.

3. A Machine That Ran — the Logic Theorist

In the middle of the summer's discussions, one event changed the air.
Newell and Simon brought not a mere conception, but a program that already ran.
Its name: the Logic Theorist.
It was a joint work with Cliff Shaw, a programmer at RAND. The two explained that it could prove, by itself, theorems set down in Russell and Whitehead's Principia Mathematica. Like a human mathematician, it fixed its goal, chose the axioms and known theorems that looked useful, and searched for a path to the proof.
It was not, of course, human intelligence itself. But neither was it mere calculation. It did not try everything by brute force; it gave priority to the paths that looked promising. It wandered — but not blindly.
There lay the prototype of what later AI research would come to call “search” and “heuristics.”
Those in the room were made to understand.
Artificial intelligence was no longer a dream made of words alone.
Small, frail, and limited in form though it was, it had already begun to run.
In the world of research there is an eternal truth: between those who argue and those who bring something that runs, the latter always win. Between a beautiful conception argued on paper and a small implementation that actually works, the latter weighs far more. This, Simon and Newell taught everyone, in the Dartmouth summer.
The Logic Theorist would later be spoken of as a representative of the very earliest AI programs.
Simon would go on to build achievements across the borderlands of psychology, economics, and computer science, and in 1978 he would receive the Nobel Prize in economics — the first person in a field adjacent to AI research to be so honored.
Newell would spend his life in pursuit of a mathematical model of human cognition. The cognitive architecture he left behind, called Soar, became one of the important legacies of AI research.
What the two of them carried into the Dartmouth summer was not a finished thing.
It was a lamp.
A small flame, easily shaken, that a wind might put out.
But on a dark sea, it is sometimes that small flame that tells a ship its heading.

4. The Moment a Name Becomes a Flag

The greatest invention of the Dartmouth summer may have been not a program, but a name.
And yet that name was not born suddenly on the lawns of 1956.
Artificial Intelligence.
An intelligence made by human hands.
The words had already been written in the proposal of the year before.
The moment of naming cannot be cut out as a single scene. At which desk, at what hour, did McCarthy first set those two words side by side? With whom did he speak, and what was said? The sources do not survive like film.
But in history there is the moment a word is born, and the moment that word becomes a flag.
The term “artificial intelligence,” set down on paper in 1955, became the flag of a community in the summer of 1956, when the researchers actually gathered.
A new science often needs a provocative name.
A name is a flag. A cause written out.
Even before there is territory, once the flag stands, people begin to gather to it. Those who gather draw the borders, make the methods, write the papers, raise the students. And in time, a name that looked empty at first begins truly to have contents.
The words “artificial intelligence” were exactly such a flag.
The riddle of intelligence was not solved in the Dartmouth summer. Nor did machines begin to think like human beings.
But a place called “AI” came into being then.
More than half a century later, those two letters would appear in the newspapers of the world, on corporate signboards, in the titles of university lectures, in policy papers, and on the small screens in people's hands.
In 1955 the name was written; in the summer of 1956 the flag was raised.
There was, as yet, no map.

5. A Ship Without Oar or Rudder

From the Dartmouth of 1956, time runs far back.
March 1771, Edo. The execution grounds at Kozukappara.
The dissection of an executed criminal was being carried out. In attendance were Sugita Genpaku, Maeno Ryōtaku, Nakagawa Jun'an, and others; they compared the plates of a Dutch book of anatomy, the Tafel Anatomie, with the human body opened before their eyes.
The plates were right.
They differed from the pictures of the body drawn in the medical books they had inherited. The organs before their eyes and the figures in the Dutch volume agreed to an astonishing degree.
Then this book must be translated.
It was natural to think so.
The trouble was that the tools for translating were almost wholly lacking.
Of the company, Maeno Ryōtaku had studied Dutch most deeply. But there was no adequate dictionary such as we have today, no ordered apparatus of translation. Genpaku and the others had to feel out the meaning word by word, set plate against text, and guess the unknown from the known.
What they had was a single Dutch book, the facts they had seen with their own eyes, and the pressing conviction that this must be carried over into Japanese.
Later, in Rangaku Kotohajime — “The Beginnings of Dutch Learning” — Sugita Genpaku set down how it had felt: like putting out upon the great ocean in a ship with neither oar nor rudder.
No oar. No rudder.
And still they put out onto the great sea.
They guessed word by word, matched each with a Chinese character, coined names for the parts of the body, and built up sentences. They did not begin because they could translate. Because they began, little by little, they became able to translate.
In 1774, the Kaitai Shinsho — the “New Book of Anatomy” — came into the world.
It became a great turning point in Japan's reception of Western medicine, and one of the books that came to stand for the flowering of Dutch learning.
The Dartmouth summer resembles it, somehow.
The researchers who gathered reached no clear conclusion. Each pressed a different methodology, and they could not agree even on what intelligence was. The participants came and went through the weeks of summer, and returned to their universities and laboratories.
Nothing was settled. Nothing was solved.
Only — they had put out from shore.
In a ship with neither oar nor rudder.

6. Three Head Temples

The true legacy of the conference revealed itself gradually, over the years that followed.
McCarthy left Dartmouth and in time moved to the Massachusetts Institute of Technology. Minsky, too, carried his research forward at MIT, where the base that would come to represent that school's AI research took shape. McCarthy then moved on to Stanford, where he founded the Stanford Artificial Intelligence Laboratory.
Newell and Simon made their seat at the Carnegie Institute of Technology — later Carnegie Mellon University. Their work left a great lineage at the crossing of AI and cognitive science.
— Stanford. MIT. Carnegie Mellon.
These three became the principal strongholds of American AI research in the years that followed.
The history of AI is not, of course, made of these three schools alone. Edinburgh, Toronto, Montreal, Bell Labs, IBM, Berkeley, and, later, the corporate laboratories. The fire was burning in many places across the world.
Still, it is certain that from the people who gathered in the Dartmouth summer, research centers and lineages of students extended outward, one after another.
Dutch learning, too, spread by way of people. After the work of Genpaku, Ryōtaku, and the rest, it was passed down through family schools and private academies, in time giving rise to places of learning such as Ogata Kōan's Tekijuku; and beyond that appeared the bearers of modern Japanese knowledge, Fukuzawa Yukichi among them.
Learning is transmitted through people.
A head temple stands; disciples go out from it; and the disciples guard the fire in other places.
In that chain, the first small seed grows into a vast forest.
The Dartmouth summer was the moment that seed was sown.

Coda: Twenty Years On

From the Dartmouth summer onward, the researchers began to speak their bold forecasts to the world.
As an emblem of that optimism, one remark is often quoted in later years. In 1965, Simon declared:
“Machines will be capable, within twenty years, of doing any work a man can do.”
McCarthy and the others shared much the same optimism.
Twenty years from the Dartmouth summer — that is, 1976.
The sea was still dark.
Machine translation had sunk into the depths of context; theorem proving struggled the moment it left its small worlds; robots stumbled on the complexity of the real world. Chess, conversation, vision, common sense — the more obvious a thing seemed to humans, the harder it proved for machines. AI research had entered a severe season of winter. Funding dried up, students drained away into other fields, and of the summer's dream almost nothing had yet come true.
Newell and Simon, Minsky and McCarthy — all of them were still living. And all of them felt, keenly, how reckless the optimism they had voiced that summer had been.
But not one of them stopped.
The ship with neither oar nor rudder was still in the middle of the great sea. The shore was not yet in sight.
Still, one rows. One goes on rowing. That was what they had decided.
The view seen by those who rowed on — we, half a century later, would come to know it better than they themselves ever could.
End of Chapter Two

Chapter Three: Wings of Wax

A Portrait of Frank Rosenblatt (1928–1971)

1. The Boys of the Bronx

The 1940s, the Bronx, New York City.
Near the last stop of the subway line stood a high school: the Bronx High School of Science. A school that gathered boys and girls selected from across the city, with an almost unnatural aptitude for mathematics and science.
In that building were two boys.
One was Frank Rosenblatt. A boy with sharp eyes, carrying a kind of heat. Once something possessed him, he would talk until his listener tired. He had a way of using not logic alone but gesture and the rise and fall of his voice, trying to make others see the future he was seeing.
The other was Marvin Minsky. Quick-spoken, his ideas leaping. But the leaps landed, often, at a distance the people around him could not follow — and landed exactly.
How closely the two spoke, no detailed record remains. They may have walked the same corridors. They may have heard the same teachers' voices. They may have reached for the same shelf in the library.
Only one thing is certain.
The two who would later split the history of artificial intelligence in two breathed, in their youth, the air of the same school.
Rosenblatt would go toward the side that believed in machines built in the brain's likeness.
Minsky would come to stand on the side that pronounced their limits.
That future, of course, was visible to no one yet. What they had, in their teens, was only a hunger for science, and the arrogance peculiar to the young — the certainty that the world ought to be explainable.
The tragedy had not yet begun.

2. Wings of Wax at Cornell

1957, near Buffalo, New York.
In a room of the Cornell Aeronautical Laboratory, Rosenblatt sat at his desk. Twenty-eight years old. He had taken his doctorate in psychology at Cornell, and was deeply drawn to perception and the workings of the brain.
Spread on the desk were paper and pencil. In the figures, circles and arrows, many of them.
It was the nerve cell of the human brain, recast in the form of mathematics.
A single cell receives several input signals, multiplies each by a “weight,” adds them together, and — if the sum crosses a certain threshold — sends a signal on to the next cell. The simplest possible mechanism.
Even a mechanism as bare as this, he believed, could “learn” something, if only the weights could be adjusted rightly. Recognize patterns of light. Tell letters apart. Distinguish sounds — such things ought, in principle, to be possible.
Rosenblatt gave the mechanism a name.
The Perceptron.
A coinage drawn from the Latin — that which perceives.
It was still a rough pair of wings. Few feathers, a simple frame, far too frail to fly the whole of the sky. But wings they were, all the same.
In those days, most researchers of artificial intelligence were trying to express human reasoning in symbols and rules. Logic, search, theorem proving. Intelligence was thought to be the manipulation of rules made explicit.
Rosenblatt's intuition was different.
Intelligence is not a thing given as rules from the start.
Is it not a thing that takes shape within, as one touches the world, errs, and is corrected?
He felt that he was assembling, at this moment, the wings on which he would rise into the sky. They had a beautiful structure, like the wings of wax and feathers that Daedalus built for his son Icarus.
Wings of wax, once assembled, really do fly.
The trouble was that no one knew how far it was safe to fly them.

3. The Newspaper's Prophecy

July 1958.
The Perceptron left the laboratory and became a newspaper headline.
The New York Times reported, at length, that the U.S. Navy had unveiled the embryo of a new computer that would learn by itself and grow the wiser for it.
The article took the future in one bold stride.
In time, machines of this kind might walk, talk, see, write, reproduce themselves — might even become conscious of their own existence.
What matters is that this grand vision should not be read as Rosenblatt's own verbatim prophecy. Into it were mixed the Navy's expectations, the newspaper's fever, and the dreams the age directed at science and technology.
Even so, it is certain that Rosenblatt himself believed boldly in the Perceptron's possibilities.
What it could do at the time was exceedingly limited.
Distinguish simple patterns.
Be given examples, and change its weights.
That was all.
But a machine whose every rule had, until then, been written in by human hands was now changing its own insides from examples.
That single point was new enough.
The newspaper saw, beyond the rough wings, a sky that did not yet exist.
Rosenblatt, too, was looking at that sky.
He did not hide the wings.
He raised them high.
Behold, he said — the machine learns.
Daedalus, it is said, told his son Icarus:
Do not fly too high. Near the sun, the wax will melt.
Do not fly too low. Near the sea, the feathers will grow wet.
But Icarus grew drunk on flight itself.
It is hard to tell one who has learned that the sky exists to keep to a middling height.
Rosenblatt flew, too.
And his flight was far too early for its age.

4. The Mark I

The end of the 1950s, Buffalo, New York. The Cornell Aeronautical Laboratory.
Rosenblatt's conception was no longer a thing of paper alone.
In the room, a device was taking shape: a receptive surface of four hundred photocells in a twenty-by-twenty array, a mass of wiring, banks of variable resistors.
Its name: the Mark I Perceptron.
It was the first full neural-network machine of the earliest era. The working machine would be shown before the press in 1960.
Seen with today's eyes, it was a primitive device.
But at the time, it looked like the future itself.
The machine could learn, from examples, the difference between simple visual patterns. No human wrote in the rules of judgment one by one. As input was laid upon answer, the weights within it changed.
A learning machine.
The words had taken on the weight of a real thing.
The United States Navy supported the research.
Newspaper reporters gathered.
Young students gathered in the laboratory.
Rosenblatt, it is said, was a devoted teacher. He listened to his students, argued without counting the hours, and spoke of research as if sharing out his own heat.
He was brilliant, at times extravagant — and he drew people to him.
And apart from the laboratory, he had one other place.
The sea.
Rosenblatt loved his sailboat.
To raise the sail, to read the wind, to glide across the water.
When the wind was right, the boat ran with astonishing lightness. The feeling of moving not by one's own strength but on an unseen current — perhaps it set him free.
A sail, too, in its way, resembles wings of wax.
Given the wind, it can fly any distance.
But when the storm comes, it falls at once into the sea.

5. The Book of Limits

1969, Cambridge, Massachusetts.
Marvin Minsky and Seymour Papert of MIT published a book.
Its title: Perceptrons.
It was not an accusatory tract, condemning the perceptron out of emotion.
It was, rather, a book that examined exhaustively, in the language of mathematics, what a certain class of perceptrons could and could not do.
Among the limits it demonstrated, one example became especially famous in later times.
Exclusive or.
XOR.
Of two inputs, true if exactly one is true. False if both are true, or both are false.
This arrangement cannot be divided in two by a single straight line.
A single-layer linear perceptron divides the world by one straight line — or, in higher dimensions, by a hyperplane. And so it cannot express problems that, like XOR, are not linearly separable.
The mathematical point was correct.
But history does not move by mathematics alone.
It was known that with more layers, the range of expressible problems grows. The trouble was that at the time, an efficient method for training such multilayer networks, sufficient computing power, and sufficient data were not yet at hand.
And so Perceptrons came to be spoken of, in later years, as the book that stood for the stagnation of neural-network research.
But a winter cannot be laid at the door of a single book.
Inflated expectations.
Meager computing power.
Limited experimental results.
The redirection of research funding.
The rise of other methods in AI.
Many causes lay one upon another.
History sometimes wants a single villain.
But a real winter has no single culprit.
Even so, it is certain that Rosenblatt's wings could not keep flying in that age.
He was still in his early forties.

6. The Sea on the Eleventh of July

11 July 1971.
Chesapeake Bay, Maryland.
It was Rosenblatt's forty-third birthday.
He went out on the water in his boat. It should have been a calm day. But wind changes. Water roughens. A boat heels at an angle no one expects.
The boat capsized.
The precise circumstances of the accident have not come down to us.
Rosenblatt did not come back.
Born on the eleventh of July; on the eleventh of July, he lost his life at sea.
One can set the coincidence aside as mere chance. As a matter of fact, perhaps one should. His death is officially recorded as an accident. That must be written plainly.
But there are moments in history when, beneath the facts, a thing takes the shape of myth.
A man who aimed at the sky on wings of wax falls into the sea and dies.
Here the story of Icarus quietly casts its shadow. Did Rosenblatt fly too near the sun — or did the sun of his age burn his wings far too early? There is no knowing.
Was he wrong?
In one sense, he was. The Perceptron did not soon walk or talk, and it held no consciousness.
But in another sense, he was right.
Machines learn from experience. They look at the world and change their weights. Not from explicit rules, but from countless examples, they take shape within.
That intuition did not die with him. It only withdrew from the stage, and sank to the bottom of a cold sea.

Coda: Half a Century On

October 1986. The journal Nature.
A paper appeared, by three authors: David Rumelhart, Geoffrey Hinton, Ronald Williams.
Its title: “Learning Representations by Back-propagating Errors.”
Send the error back, from the output side toward the input side, and correct the weight of each connection, little by little.
Backpropagation.
The idea itself had its forerunners. But the 1986 paper showed, vividly, that multilayer neural networks could learn internal representations, and it pushed the method back to the center of neural-network research.
The XOR that the single-layer perceptron could not cross, a multilayer network could.
The road toward more complex problems, too, began to come into view.
Not that winter ended overnight.
Computers were still weak.
Data was still scarce.
The true dawn would need more time yet.
And then, 2012 — AlexNet.
And then again, 2022 — ChatGPT.
Of the machines the 1958 newspaper had dreamed — machines that see, learn, write, and talk — some became real, through the joining of technologies that had grown up separately.
Self-reproduction and self-consciousness remain questions to be handled with care.
But at the least, a world arrived in which machines learn from examples, tell images apart, write prose, and converse with human beings.
Rosenblatt's picture of the future was not right in everything.
Only, some of its directions had been visible to him half a century early.
The wings of wax melted once.
They fell into the sea.
People told it as a story of failure.
But a later age rebuilt the same dream from other materials.
Stronger computers.
Larger data.
Deeper layers.
Subtler methods of learning.
The wings that had been made of wax and feathers were reassembled from silicon, mathematics, and electric power.
If Rosenblatt could see the world of half a century on, what would he say?
Probably he would laugh, a little extravagantly.
You see? — he would say.
The sky I was looking at — it was there after all.
The wings that fell into the sea on the eleventh of July had not vanished.
They had only lain sunken, a long while.
And in the twenty-first century, those wings once made of wax began, in another material, to beat again — inside the machines of all the world.
End of Chapter Three

Chapter Four: The Hidden Flame

Winter and the Scholars of Dutch Learning (1969–1986)

1. The Coming of Winter

1973, Britain.
A single report sent a cold wind through the world of AI research.
Its author was Sir James Lighthill — an authority on fluid dynamics, not a man who had led research as a specialist in AI. The British Science Research Council had asked him to evaluate the state of artificial intelligence research.
Lighthill's judgment was severe.
Its purport: measured against the great promises of its early days, AI research had not delivered results enough.
After this report, public support for AI research in Britain was drastically reduced.
In the United States too, ARPA's funding came, by degrees, to demand clearer practical results.
Seventeen years from the Dartmouth summer.
Much of the optimism the early researchers had voiced had passed its deadline unfulfilled.
Machine translation had struck the wall of context.
Image recognition suffered before the complexity of the real world.
Speech recognition, too, grew hard the moment it left its narrow conditions.
Neural-network research, amid the limits of the simple perceptron, the shortage of computing power, and the shifting currents of the field, had drifted far from the mainstream.
The first winter of AI had arrived.
But winter does not come in a single day.
One report, one book, one failure does not freeze a world.
There are inflated hopes, and disappointments, and the redirection of money, and the migration of researchers' interests.
When these pile one upon another, the season turns.
The Dartmouth summer had grown distant.

2. A Neo-Confucian Spring

And yet within the winter, in another place, an unexpected spring had come.
Stanford University, California.
The Nobel laureate Joshua Lederberg and the computer scientist Edward Feigenbaum, with others, were engaged in a bold attempt.
To move human expert knowledge into the machine.
Not to build a universal intelligence all at once.
Choose a narrow domain, and put the knowledge and the lines of judgment its experts use into a form a machine can wield.
The first representative attempt was made in the domain of organic chemistry.
The program's name was DENDRAL.
Taking data such as mass spectra for its clues, it inferred the candidate structures of organic compounds.
There, the specialist knowledge of chemists was joined to a mechanism of search.
It was not a universal intelligence.
It could not converse, knew nothing of the world's common sense, could write no poems.
But in its narrow domain, it was strong.
DENDRAL gave AI a lesson.
Rather than dream of an intelligence too broad, might not a machine be more useful given deep knowledge in a narrow field?
Feigenbaum pressed this way of thinking outward under the name of “knowledge engineering.”
In the metaphor of this book, it resembles, a little, the Neo-Confucianism of the Tokugawa age.
The world has an order, and that order can be put into words.
Arrange the knowledge to be learned, transmit it, apply it in the fitting place.
As young men in the domain schools and private academies sought to learn the order of the world through the words of the classics, symbolic AI sought to hand human knowledge to the machine in explicit form.
Not, of course, that Neo-Confucianism and AI were the same thought.
What resembles is the gesture: to describe knowledge, to systematize it, to give it a form that can be inherited.
That conviction became, from the 1970s into the first half of the 1980s, a great current of artificial intelligence research.

3. MYCIN in the Clinic

The Stanford School of Medicine. A young physician took note of Feigenbaum's method. Edward Shortliffe. At the time, in his late twenties.
Shortliffe thought to apply the method to one specialist domain.
The diagnosis of blood infections, and the choice of antibiotics.
In the hospitals of that day, this was a difficult territory where judgments often divided. Diagnoses differed from physician to physician, and so did the drugs prescribed. It was a succession of delicate judgments on which a patient's life could turn, and in which experience spoke.
Shortliffe went from specialist to specialist, and drew out, with care, the flow of their judgments.
“If bacteria of this kind are detected in the patient's blood, and the patient's age falls within this range, prescribe antibiotic X.”
“If the patient has a history of this kind, and this value in the blood test runs high, give antibiotic Y priority.”
— Rules of this sort he gathered, some five hundred of them.
He built them into a machine.
The machine's name was MYCIN. MYCIN was not a machine that ran on a bare “yes” or “no.” To handle uncertainty, it took in the idea of certainty factors. As a physician's judgment is not always a hundred-percent verdict, so the machine, too, tried to handle the shadings of possibility.
In the late 1970s, early evaluation experiments were carried out. MYCIN's judgments were set against those of the specialists. In those evaluations, MYCIN showed results that stood comparison with the experts.
This was a shock.
A machine had drawn near to the judgment of physicians.
Not that MYCIN came to be widely used in hospitals as it stood. The seat of responsibility, the labor of input, the joints with the medical system, the trust of the ward. Real medicine had walls that diagnostic accuracy alone could not cross.
Even so, MYCIN's meaning was large.
If a machine holds an expert's knowledge as rules, it may answer real problems. That expectation spread into industry.
Around the same time, at DEC — the Digital Equipment Corporation — an expert system called XCON was put to work automating the configuration of computers, and was bringing savings of tens of millions of dollars a year.
Artificial intelligence began to be spoken of, once more, in the language of use.
Feigenbaum said it:
Knowledge is power.
It became the watchword of the age of expert systems.
The Neo-Confucian spring seemed, now, in full bloom.

4. The Dream of the Fifth Generation

That spring crossed the Pacific and reached Japan.
1982, Tokyo.
In April 1982, under the lead of the Ministry of International Trade and Industry — MITI — a national project was set in motion.
The Fifth Generation Computer Project.
The name came from a division of the computer's history into five generations. The first generation of vacuum tubes; the second, of transistors; the third, of integrated circuits; the fourth, of very-large-scale integration — and the fifth, endowed with artificial intelligence.
It was a grand national design: Japan would overtake the world's computer industry at a stroke, by means of AI.
The budget: some fifty-four billion yen over ten years. For its day, a scale to make America and Europe tremble.
The organization at its head was the Institute for New Generation Computer Technology — ICOT, for short. Its research base was set in Mita, Tokyo.
At the center of the project stood Kazuhiro Fuchi, forty-six years old — a specialist in logic programming, seconded from the Electrotechnical Laboratory.
Fuchi's work was a researcher's work, and at the same time a translator's. To carry the researchers' words to the government; to carry the government's design to the companies; to turn the engineers of the companies toward one great goal. From all over Japan, young computer scientists, corporate researchers, and engineers gathered to ICOT.
There was real heat there, in those days.
Japan would lead the next age of the computer.
Not a machine for mere numerical calculation, but a machine that handles knowledge.
To realize a computer that reasons by logic, handles natural language, and answers like an expert.
Abroad, there was alarm. America was alarmed. Under the Reagan administration, the Strategic Computing Initiative was launched; in Europe, rival projects — Esprit, Alvey — rose one after another. Japan's challenge was taken that seriously. The core technology ICOT chose was logic programming.
To develop the logic languages represented by Prolog, born in Europe in 1972, and run them at scale upon parallel inference machines. Humans would not write the details of procedure; give the machine facts and rules, and it would draw the answers by inference.
This, too, stood upon a Neo-Confucian view of the world.
The world can be described by logic, and reasoned over by logic.
Kazuhiro Fuchi believed in that ideal more than anyone.

5. A Forest of Rules

But spring does not last.
In the late 1980s, the limits of the expert systems began to show.
At first it had seemed that the more rules you added, the wiser the machine would grow. In practice, the more they grew, the more they tangled. Expert A's rules collided with expert B's. Put in new knowledge, and part of the old rules broke. Add an exception, and the exception needed exceptions of its own.
Knowledge did not stack itself in order, like books upon a shelf.
The greatest trouble of all was common sense.
The knowledge of a specialist domain can still be drawn out. Ask the physician, the chemist, the engineer. But the common sense human beings use in daily life was too broad, and lay too deep in the unspoken.
“Water, if spilled, falls downward.” “A person who does not eat grows hungry.” “When night comes, the sun sets and the moon rises.” — There was research that tried to write out such things, one by one, in the form of rules. But to write out all the common sense of the world was a labor near to infinite. The more obvious a thing was to a human being, the harder it was to set down as a rule.
The expert systems were wise in narrow rooms. Step into the corridor, and they lost their way. Outside the building, they knew almost nothing.
In time, the market for dedicated LISP machines began to crumble as well. The expensive special-purpose machines were run down by general-purpose workstations whose performance climbed at speed. The companies grew cautious, once more, of the words “artificial intelligence.”
The Fifth Generation project, likewise, could not realize its first grand goals as they stood.
Run this story's clock a little forward: in 1992, its ten-year term complete, ICOT came to a close of accounts. Parallel inference machines, logic-programming systems, research in knowledge processing — there were results, certainly. But the “intelligent fifth-generation computer” had not transformed society.
At the evaluation meeting, Kazuhiro Fuchi reported the project's results in a level voice.
He did not step aside from his responsibility.
Measured against the first grand design, the achievement had stopped at a part of it. That judgment he accepted, facing it straight on.
But there was something he pointed to, up to the very end.
The conviction that the engineers who had grown here would carry the information technology of the Japan to come.
Later, that conviction became fact. The researchers who came out of ICOT scattered into the core of Japanese computer science, and in each place continued their long contributions.
Neo-Confucianism, even as it lost its own force, had furnished the people who bore the Meiji reforms.
For the progress of AI, however, a second winter had arrived.

6. The Lamps of the Dutch Scholars

In the winter, in another place, another flame went on quietly burning.
Edinburgh, Scotland.
In a certain laboratory there was a young psychologist. Geoffrey Hinton. Born in London, 1947. At the University of Edinburgh, he was working to reproduce human cognition inside the computer.
He went on stubbornly believing in neural networks — that field Rosenblatt had left behind, the one now labeled “out of date.”
Born in London, schooled in artificial intelligence at Edinburgh, he was strongly drawn to neural networks and distributed representations. Knowledge is not necessarily written down as symbols, one by one. Might it not dwell, distributed, among a multitude of weights?
Seen from the orthodoxy of symbolism, this was heresy.
Meaning is not a thing made explicit like the entries of a dictionary.
Rules are not a thing human beings write out in full.
The machine, through examples, changes its own inner form.
Rosenblatt's fire had not yet gone out.
Hinton changed posts several times. From Edinburgh to the University of California, San Diego — where he met David Rumelhart and the young researchers called the PDP group. Then to Carnegie Mellon. At last to the University of Toronto, in Canada.
Of his reason for moving to Toronto, he said in later years, half in jest:
“In America, funding for AI research must keep military uses in view. I disliked that, so I came to Canada, where the military color is faint.”
Around the same time, David Rumelhart and others in California were rearing the school of thought called parallel distributed processing — PDP. A view that takes human cognition not as a central command tower, but as the interaction of a multitude of simple units.
In France, Yann LeCun was at work on neural networks. He would later apply convolutional neural networks to the recognition of handwritten digits, and leave a great footprint on the stream of image recognition.
And of a still younger generation, Yoshua Bengio of Canada followed.
The three would later be called the three great figures of connectionism.
But in the 1980s, in AI's second winter, they stood in a place plainly apart from the mainstream of the research world.
The three read one another's papers, and sometimes met at the same conferences. Between them ran a loose fellowship.
“The neural network is not finished yet.”
That was the three men's modest faith.
The flames that move history often begin at just that size.

Coda: 1986, the End of Winter

October 1986.
In the journal Nature, a paper appeared.
By David Rumelhart, Geoffrey Hinton, and Ronald Williams: “Learning Representations by Back-propagating Errors.”
The idea called backpropagation had its forerunners from before.
But this paper showed its power, vividly, as a method for training multilayer neural networks.
The output errs.
Send that error back, from the later layers toward the earlier.
Compute how much each weight contributed to the error, and correct it, little by little.
Repeat, over and over.
By this mechanism, the road to actually training multilayer networks opened wide.
The XOR that the single-layer perceptron could not cross became, now, a small exercise.
The door toward more complex problems had opened, by a crack.
Not, of course, that winter ended with this.
Computers were still weak.
Data was still wanting.
Neural networks drew notice again, but to change the world they would need much more time.
The true dawn would have to wait another twenty-six years.
2012.
AlexNet.
But already, in 1986, beneath the snow there was the breath of spring.
As the great cathedral of Neo-Confucianism swayed and the words of orthodoxy began to lose their force, the scholars of Dutch learning were guarding their small lamps.
They were not yet victors.
The world did not yet know what those lamps meant.
Even so, the fire had not gone out.
That flame, burning on in secret in the depths of AI's winter, would in time light the world of the twenty-first century.
End of Chapter Four

Chapter Five: Black Ships and Samurai

The Statistical Revolution and Deep Blue (1990–1997)

1. The Eve of the Opening

At the opening of the 1990s, artificial intelligence research lay once more in deep fog.
The spring of the expert systems had passed. Write knowledge in as rules and the machine grows wise — that faith had lost its force before the bottomless marsh called common sense.
Japan's Fifth Generation Computer Project, too, for all the grandeur of its ideals, had not been able to realize the future as promised. The dream of logic programming and inference machines left solid technology and trained people behind it, but did not repaint the age entire.
The neural network had seemed to draw breath again with the backpropagation work of 1986. But it was not yet the mainstream. Computers were weak, data was scarce, and the world met the words “neural network” with half-belief.
Outside the signboard of artificial intelligence, other fields of research were gathering strength.
Machine learning.
Pattern recognition.
Statistical estimation.
Data analysis.
These were no mere aliases of AI. Each was a discipline with its own history and methods.
But the result was that even in an age when AI could hardly speak its great dreams aloud, around its edges, the technologies that would make the next age went on growing.
The air of it resembles, a little, Japan at the end of the shogunate.
The Tokugawa world had lasted long. There was order, there were rules, there were ranks, there were words. But in the middle of the nineteenth century, from outside that order, another force was drawing near. Everyone felt that something would change. What it would change into, no one yet could see.
AI research was the same.
The shogunate of symbol and logic had not yet fallen.
But offshore, black shapes could already be seen.
Those shapes were not steamships.
They were black ships of knowledge, made of equations and data.

2. The Black Ships Came as Equations

The third day of the sixth month, Kaei 6. By the Western calendar, July 1853. Off Uraga.
Four foreign ships appeared. Hulls painted black. Black smoke rising from their stacks. Ships that moved without wind — ships that crossed the sea not by sail, but by steam.
Commodore Perry carried a letter from the American President, Fillmore.
Japan should open its doors, and have commerce with the United States.
The shogunate shook.
Open the country, or expel the barbarians? Defend, or change?
The black ships did not conquer Japan on the spot. But they shattered the view of the world that had held until then. Japan could no longer close the world within its own interior. That fact, the ships trailing black smoke set before every eye.
In the 1990s, a similar set of “black ships” had appeared in AI research.
They were not one ship come from one country.
They were several squadrons, risen from seas different from one another.
Vladimir Vapnik studied the boundaries of statistical learning.
Judea Pearl handled uncertainty with webs of probability, and later pressed causation itself into the reach of computation.
Frederick Jelinek handled speech and language with great volumes of data and models of probability.
The three were not comrades under one banner.
Their subjects, their methods, their philosophies differed.
But watch their work from offshore of the same era, and a single change comes into view.
One need not rewrite the whole world, by human hand, into rules of “if A, then B.”
Uncertainty can be handled as probability.
The boundary of a classification can be learned from data.
Even the sequences of words hold statistical regularities.
This meant that another way of making intelligence — not symbolism alone — had begun to hold power.
The steamship did not erase the sailing ship in a single night.
Nor did statistics erase logic in a single night.
But the way ships crossed the sea had begun to change.

3. Boundary Lines and Webs of Probability

Vladimir Vapnik was a mathematician who had long studied the theory of statistical learning in the Soviet Union.
His work grew up on the far side of the Iron Curtain. When in time he moved to the United States and joined Bell Labs, the theory drew great notice in the research world of the West as well.
The support vector machine. SVM.
It was a powerful method for classifying data.
Not merely to draw a boundary line.
To seek the boundary that divides two populations with the largest possible margin.
To seek a division that does not cling too tightly to the training data, and can bear unknown data as well.
The machine does not memorize rules.
From data, it learns a good boundary.
Judea Pearl, meanwhile, had his eyes on the world of uncertainty.
Reality is not made only of true and false.
A patient's symptoms may raise the likelihood of one disease, while the possibility of another remains.
From the events observed, the unseen causes must be inferred.
The Bayesian network Pearl systematized expressed the dependencies among variables as a web of probability, and became a powerful instrument for reasoning under uncertainty.
And his work pressed further, toward the explicit consideration of causation itself.
For artificial intelligence, this was a great turning.
Most AI until then had loved the world of logic.
Correct, or incorrect.
Provable, or not.
But real intelligence is muddier.
Human beings judge, day after day, on incomplete information.
Then machines, too, must be able to handle uncertainty.
SVM and the Bayesian network are not the same theory.
Yet each, from its own direction, widened AI's center of gravity.
Not logic alone, but learning.
Not certainty alone, but probability.
The opening of a civilization is not merely the arrival of new tools.
It is the increase of ways of thinking themselves.

4. Every Time I Fire a Linguist

One of the emblematic battlefields of the statistical revolution was speech recognition. Frederick Jelinek, a researcher at IBM, was a Czech-born researcher of Jewish descent. Bearing the shadows of war and exile, he crossed to America, and became in time the central figure of the research that treated speech and language by statistics.
Speech recognition was a hard problem for AI.
The sounds human beings speak are ambiguous; they collapse, break off, overlap. The same word is pronounced differently by different speakers. Noise enters. Spoken language does not proceed by the grammar book.
On the older way of thinking, linguists should set the rules of grammar in order, write the rules of pronunciation, and teach them to the machine. To understand language, one must know the structure of language explicitly — so it was held.
Jelinek doubted that thought.
Must a machine really understand grammar?
Show it great volumes of speech set against text, and might it not learn, statistically, which sounds tend to answer to which words, and which word tends to follow which?
The hidden Markov model.
The N-gram.
Neither the hidden Markov model nor the N-gram understands the meaning of words as a human does. Taking the words and states that have appeared so far for its clues, it judges by probability what is most natural to come next. That is: a mechanism that chooses not meaning itself, but the likeliest sequence. It treats language not as a system of rules but as a current of probabilities. With such probabilistic models, a phenomenon can be described and predicted in the form: what is the probability that this outcome occurs?
He is reported to have said, in a tone between jest and earnest:
“Every time I fire a linguist, the performance of the speech recognizer goes up.”
— To give up the rules of language is to advance the recognition of language.
It was a sharp letter of challenge to symbolic AI, founded as it was upon logic.
As the Meiji government dismantled the institutions of the samurai with the sword-abolition edict and the abolition of the domains, the statistical revolution lowered handwritten rules and expert intuition from the center of AI.
Of course, the vanishing of the samurai did not make the art of the sword, or its ethics, meaningless on the instant. In the same way, neither linguistics nor logic vanished.
Only, at the center of the new age, something else took the seat.
Data.

5. The Last Samurai of Chess

The time: the late 1990s.
The world chess champion, Garry Kasparov.
Born 1963 in Baku, in Soviet Azerbaijan. Soviet junior champion at twelve; world junior champion at seventeen; at twenty-two, the youngest world champion of his day — a man everyone acknowledged as one of the strongest chess players in history.
He would hold the world's number-one ranking into the 2000s. For more than ten years, an absolute king, broken by no one.
Kasparov was, in a certain sense, like Saigō Takamori.
Saigō, though a chief architect of the Meiji Restoration, fought the new government's army in the end, was defeated, and ended his own life. He was a man who embodied to the last the values of the traditional samurai — honor, loyalty, the clean death. He had seen the coming of the new age earlier than anyone and prepared its arrival; yet he himself could not settle into that new age, and fell together with the old world.
Kasparov was somewhat different — but the structure was alike.
He was the highest peak of human chess. At the same time, he belonged, within the history of knowledge, to a certain last generation — the last generation for whom it was possible that the human reads deeper than the machine.
And against him, a machine appeared, offering battle.
Kasparov, too, was no mere loser of an old age. He did not slight the progress of machine chess. He used computers in his study, and took them into his preparation. He watched the new age from closer than anyone.
But at the final line, he believed.
Human intuition surpasses the machine's calculation. The deep reading of a world champion does not fall to brute-force search.
And against that belief of his, a machine appeared, offering battle. IBM's Deep Blue.
A calculating machine built for chess alone. A beast of computation that read two hundred million positions in a second.
Kasparov, at first, took it lightly.
“A machine cannot beat human intuition” — so he believed.

6. New York, Nineteen Moves

February 1996, Philadelphia.
The first official match between Kasparov and Deep Blue was played.
Game one: Deep Blue won.
A world champion had lost a game to a computer at standard time controls. It was a historic event.
But Kasparov did not crumble. From the second game on, he probed the machine's weaknesses, applied pressure, shook it. The final result: three wins, one loss, two draws for Kasparov. Four points to two.
The human had still won.
The world breathed out.
The machine is strong. But it does not yet reach the strongest human.
One year later.
May 1997, New York.
The rematch was set at the Equitable Center. Deep Blue had been improved. Faster, deeper, more artfully tuned.
Game one went to Kasparov.
But in game two, the air changed.
Deep Blue played a move beyond Kasparov's expectation. Not a mere grab of material, not a short-term gain. It was a move that seemed to carry a long-range conception. To Kasparov, it looked like a human move.
He lost the game.
And a doubt took hold of him.
Is this machine truly playing on its own?
Is some human master intervening, somewhere? The IBM side denied it. There is no evidence of improper human intervention. But what mattered was not only the presence or absence of proof. That doubt had been born in Kasparov's mind — that in itself weighed heavily.
The most dangerous thing on the board is not the opponent's move.
It is the wavering of one's own mind.
Games three, four, and five were drawn.
And then, game six.
Kasparov chose the Caro-Kann Defense. A solid, cautious opening. But that day he lacked his usual attacking spirit. Entering safety, he invited danger instead.
In the opening, he left a fatal gap.
Deep Blue did not miss it.
Nineteen moves.
In a mere nineteen moves, Kasparov resigned.
The world chess champion had lost an official multi-game match to a machine.
In that moment, not only the history of chess but the history of artificial intelligence changed.
In the press room, Kasparov was worn through. There must have been anger, and doubt, and humiliation. But what was present there was not only the defeat of a single player.
A symbol of human intellect had gone down upon its knees before a machine.
So the world received it.

Coda: After the Black Ships

Deep Blue's victory does not mean that artificial intelligence thought as a human does.
That is an important point.
Deep Blue did not love chess as humans love it. It did not tremble at a beautiful move, nor fear defeat. It simply read an enormous number of positions, evaluated them, and chose the move that seemed best.
And with that, it won.
That “with that” is what moved history.
Perry's black ships did not occupy Japan on the spot, either. But they showed that the world's measure had changed. The inner order alone would no longer suffice. The force come from outside had to be acknowledged, taken in, remade.
Deep Blue was the same.
Human intuition alone was no longer enough.
One could no longer speak of intelligence while ignoring the machine's power of calculation. AI research was turning into a thing that actually runs, is measured, and wins competitions.
In the same season, the statistical revolution was advancing in the laboratories. Speech recognition, character recognition, machine translation, information retrieval. In every field, the methods that learn from great volumes of data began to overpower the handwritten rules.
The coming of the black ships led to the Meiji Restoration.
And in AI as well, the shock of the black ships brought forth a new order.
The shogunate of logic alone was ended; the age of data, probability, and computing power began.
Kasparov went on criticizing the Deep Blue match for long afterward. Something in it he could not accept. It was the anger of the defeated, and the pride of a king.
But he did not end there.
In time he proposed a new chess in which human and machine fight as one. Advanced chess — or freestyle chess. Not human against machine. A human and a machine, joined, against another pair of human and machine.
What came to light there was an interesting fact.
In the early tournaments at least, it was not the strongest human who won alone, nor the strongest machine that won as it stood. The winner was the human who used the machine best. Human judgment and machine calculation. Intuition and search. Experience and data. That combination was strongest of all.
Out of defeat, Kasparov had found another answer.
The machine is not the end of the human.
It is the human's new instrument.
It resembles the Japan of Meiji, which took in the technologies of Europe and America and yet did not end in mere imitation, but tried to make a modernity of its own.
The first emotion of an opened country is humiliation.
But history does not end at humiliation.
Beyond it comes the age of learning, of translating, of mingling and remaking.
The statistical revolution and Deep Blue brought that age to AI.
And as the twenty-first century opened, data grew greater still, computers faster still, the methods of learning deeper still.
2012. AlexNet.
That dawn was already drawing near, from far off.
End of Chapter Five

Chapter Six: The Workshop at Dawn

The Deep Learning Revolution and AlexNet (2006–2012)

1. The First Sign of Spring

2006, the University of Toronto, Canada.
Geoffrey Hinton had turned fifty-eight.
For a long time, he had gone on guarding neural-network research. Edinburgh, California, Carnegie Mellon, and then Toronto. He changed countries, changed universities; the fashions of the age changed; his conviction did not.
Intelligence is not necessarily written down as explicit rules.
Knowledge is not a thing set in order upon the shelves of symbols.
It dwells, distributed, in a multitude of connections.
But the world long refused to believe it.
The neural network was looked on as an old dream. Rosenblatt's Perceptron was a wing that had flown too near the sun nearly half a century before, and fallen into the sea. Hinton went on with his research as if gathering up the fragments of that sunken wing.
Unlike Rosenblatt, who spoke of the future as if from a stage, Hinton was closer to a craftsman of the workshop.
A man who fed the furnace each morning, so that the fire would not go out.
And in 2006, that man showed one important road.
Neural networks with deep layers were, at the time, hard to train.
The problem of weakening gradients.
The difficulty of optimization.
The shortage of computing power and data.
Even knowing that deeper was better, the methods for teaching such depth were wanting.
Hinton and his colleagues showed a method of pre-training the layers one at a time and stacking them, rather than training them all at once. Deep belief networks, built with restricted Boltzmann machines.
The paper's title: “A Fast Learning Algorithm for Deep Belief Nets.”
This work called interest in deep neural networks strongly back to life.
Hinton did not invent the words “deep learning” themselves in 2006.
But from around this time, the current of research into training deep layers gathered speed, and in time the name “deep learning” became the flag of the age.
It resembled, a little, Petrarch in fourteenth-century Italy.
Petrarch sought out the writings of ancient Rome that slept in the archives, and made people remember the worth of the classics.
Hinton, perhaps, was doing the same.
The neural network is not dead.
It is only that the age has not yet caught up with it.
The long Middle Ages were coming to an end.

2. The New Paints

A renaissance is not born of a single genius.
A genius needs tools.
Needs a workshop.
Needs pigments.
Needs patrons.
And needs chance.
To deep learning, too, an unexpected tool appeared. The GPU. The GPU was not born for artificial intelligence. It was a component for drawing game screens smoothly. Light moving through three-dimensional space, shadow, objects, explosions, smoke. To compute these in an instant and put them on the screen, the GPU was built to carry out great volumes of simple calculation in parallel.
But — at some point, someone noticed.
“Is not the computation of a neural network, in its essence, a great volume of parallel calculation?”
Each nerve cell computes its signal independently. If all of it could be processed in parallel at once, the speed of training would rise by orders of magnitude.
The training of a neural network is a succession of enormous matrix calculations. Multiply countless weights, add, send the error back, change the weights again. Each single calculation is simple; the number is vast.
Then let the GPU do it.
Around 2007, NVIDIA released CUDA, opening the GPU to general-purpose computation beyond the drawing of images. Researchers began to divert it to the training of neural networks.
The result — computing speed leapt by a hundredfold and more.
Training that had taken weeks now ended in realistic time.
Networks of a size that could not be tried became triable.
Depths that had been given up came within reach.
A component polished for gamers had been carried into the workshop of artificial intelligence.
It resembled the oil paints of the Renaissance.
Until then, painting had centered on the mural and the tempera panel. But in the fifteenth century, when the Van Eyck brothers improved oil paint, painters became able to make pictures of extreme fineness that could be corrected any number of times. Leonardo da Vinci's Mona Lisa, Titian's opulent color — without the new technique, neither would have been born. The GPU was the oil paint of the age of deep learning.
On the shelf of the workshop, a new paint had been set.

3. Fei-Fei Li's Cathedral

But tools alone were not enough.
To show the world to a machine, the world itself was needed.
In the late 2000s, Fei-Fei Li of Princeton — she would later move to Stanford — was looking at the stagnation of image recognition from another angle. The problem was not the algorithms alone. Was it not that the machine had never sufficiently seen the world in the first place?
“To show the world to the machine, there are not enough pictures of the world.”
Image-recognition research until then had trained machines on a few tens of thousands of images at most. But a human child, within a few years of birth, has seen hundreds of millions of images. For a brain to learn, might not quantities of that order be necessary — so she hypothesized.
A human child grows through an enormous volume of visual experience. Dog, cat, chair, plate, car, tree, sky, face, shadow, rain, snow. The world is not so poor a thing that tens of thousands of training images could suffice for it.
Then the machine, too, must be shown the world.
Fei-Fei Li conceived a gigantic database of images. Taking the English lexical system WordNet as its foundation, she organized the categories of objects, gathered images from the internet, and had human hands attach the labels.
Cat.
Dog.
Fire engine.
Lighthouse.
Strawberry.
Chair.
Kingfisher.
Airplane.
Toaster.
Violin.
Their number came, in time, to exceed ten million.
The work was carried by nameless laborers all over the world, taking part through Amazon Mechanical Turk. They looked at the images one by one and answered what each one was. The pay was small, and their names would almost never enter history.
But without that handwork, the revolution to come would not have come. ImageNet.
The database, announced in 2009, was the great cathedral of the age of deep learning.
A medieval cathedral was not raised by one architect alone. Stonemasons, carpenters, glaziers, carriers, donors, those who prayed. Countless hands, across generations, pushed the spires into the sky. ImageNet was the same.
The conception of one researcher and the hands of nameless workers across the world built the great cathedral of AI research.
And Fei-Fei Li prepared one more device. The ImageNet Challenge.
Laboratories all over the world would bring their image-recognition systems and compete for accuracy on the same data. Who could tell the world apart best? The stage was set.
And this competition would become, in time, the stage of the revolution.

4. The Hinton Workshop

Toronto again.
In Hinton's laboratory there were two young researchers.
Alex Krizhevsky. A computer scientist born in Ukraine and raised in Canada. Taciturn, strong at implementation, able to worry the details of a computer with obsessive persistence. He was a man who built things that ran, rather than dressing up theory.
Ilya Sutskever. Born in the Soviet Union, moved to Israel as a child, then crossed to Canada. Precocious, gifted with abstract intuition, he believed deeply in the possibilities of deep learning.
Their teacher, Hinton, placed a quiet trust in them.
The laboratory resembles the workshop of Verrocchio in fifteenth-century Florence.
Verrocchio was a sculptor, a painter, the master of a workshop. Young talent gathered there. Among them was Leonardo da Vinci. The master taught the apprentice his art; the apprentice tried to pass beyond it.
In Verrocchio's Baptism of Christ there is an angel said to have been painted by the young Leonardo. The story goes that the angel was so beautiful that Verrocchio never took up the brush again.
True or not, the anecdote speaks the essence of what a workshop is.
A workshop is not a place where the master completes everything.
It is a place where the fire the master has guarded is made to blaze up by the apprentice, in another form.
Hinton's workshop was such a place.
Hinton had the conviction and the intuition of long years.
Krizhevsky had the power of implementation to set them running in reality.
Sutskever saw the theoretical horizon that opened beyond.
And at a certain moment, the three decided.
“Let us enter the ImageNet Challenge of 2012.”
At the time, the mainstream of image recognition was hand-designed features. Humans thought out where in the image to look, extracted contours, local features, gradients, textures, and passed them to a classifier. It was a world of artisan skill.
Those who seriously believed a neural network could win were still few.
But the three of the Hinton workshop were looking at something different.
Features are not a thing human beings carve out.
They are a thing the machine itself learns from data.
That conviction they poured into a single network — into deep learning.

5. The Shock of Deep Learning

Autumn 2012. The results of the ImageNet Challenge were announced.
The system submitted from the Hinton workshop was called, at the time, SuperVision. Later, people came to call it AlexNet, after its principal implementer, Alex Krizhevsky. AlexNet was a convolutional neural network of eight layers. Five convolutional layers and three fully connected. About sixty million parameters. It was trained on two NVIDIA GPUs.
In it were many devices. ReLU.
Dropout.
Data augmentation. Fast computation on GPUs.
But most important of all was that these had been assembled into one practical, gigantic system.
The result was a shock. AlexNet's top-five error rate: about 15.3 percent.
The second-place system: about 26.2 percent.
The gap exceeded ten points.
In the competitions of research, even a one-percent difference is large. A few percent is decisive. But a gap of more than ten percent was no longer an improvement along the same line.
It was a fault-line in the age.
People understood.
This is no accident.
This is no sleight-of-hand tuning.
Deep learning is real.
The neural network, so long called out of date, stood at the summit of image recognition.
Krizhevsky, they say, was as level as ever.
Sutskever sensed the great change that lay beyond this victory.
Hinton was quiet.
He had no need to exult.
For one who has crossed the winter, spring is not a thing to be proclaimed in a loud voice.
It is enough to know that the snow has melted.
Hinton's victory was a quiet victory.

6. The Renaissance Opens

With AlexNet's victory, the world changed — quite literally — overnight.
That “overnight” was no exaggeration.
October 2012 — within a few months of AlexNet's results coming out, the very language of the industry changed. Google, Microsoft, Facebook — the great companies began, all at once, to hire away neural-network researchers at high salaries. A field of research that almost no one had cared for the year before became, suddenly, the most advanced and most important technology in the world.
Hinton himself made a decision.
At the end of 2012, he put the small company he had founded, DNNresearch, up for auction.
A company — but there was no great plant in it, no products on shelves. What there was: the three minds of Hinton, Krizhevsky, and Sutskever, and the technology they carried. Google, Microsoft, Baidu, DeepMind — four companies bid.
The next year, 2013, Google acquired DNNresearch. The price was reported at about forty-four million dollars.
In three minds, the giant companies of the world had seen that much worth.
Hinton joined Google. Krizhevsky, too, stayed within that current for a time. Sutskever would later walk another road. In 2015, he joined in founding OpenAI, and in time would carry the technical core of the GPT line.
But that is a story for later.
What happened from 2012 into 2013 was no mere fashion in technology.
It was the opening of a renaissance.
As the techniques polished in the workshops of Florence spread in time to Rome, to Venice, to Flanders, the fire of the Hinton workshop passed to laboratories all over the world.
The neural network was heresy in a corner no longer.
Deep learning became the central word of AI research.
The long Middle Ages had ended.

Coda: From the Workshop to the World

From the Hinton workshop, the apprentices spread out into the world.
Krizhevsky, after a period at Google, had a season in which he kept his distance from AI.
Sutskever went to OpenAI. Later, as its chief scientist, he would carry the core of the development of the GPT line of models. Later still, he would leave OpenAI and found another research organization — but that is a story for much later.
Hinton himself continued his research at Google, and in 2018, together with Yann LeCun and Yoshua Bengio, received the Turing Award, the highest honor of computer science.
The three who had survived the winter were welcomed, at last, into the center of the age.
From Bengio's lineage, moreover, came Ian Goodfellow.
In 2014, he announced the generative adversarial network — the GAN.
By setting two networks in contest, it generates new things that resemble the data.
The GAN carried image generation by deep learning a great step forward.
Afterward, from another lineage, diffusion models rose, leading on to image-generating AI such as Stable Diffusion.
One invention does not beget the next in a straight line.
The fire of the workshop branches, and burns another way in another furnace.
The apprentice becomes a master, and hands it on to the next apprentice.
So technology spreads, as a lineage of persons.
As Leonardo came out of Verrocchio's workshop.
As modern knowledge spread out of the small rooms of the Dutch scholars.
Out of Hinton's workshop spread the AI of the twenty-first century.
Rosenblatt's wings of wax fell, once, into the sea.
Hinton and his people took up those wings and rebuilt them in other materials.
New wings, made of GPUs and data and learning algorithms.
In 2012, those wings flew at last.
And the flight did not stop.
The flame lit images first.
Then it lit speech.
In time, it would light words.
And in 2016, on the board of Go, the world would catch its breath once more at a machine's move.
That move would later be called a “divine move.”
End of Chapter Six

Chapter Seven: Two Divine Moves

AlphaGo versus Lee Sedol (2016)

1. The Four-Thousand-Year Board

Go is an old game.
Chinese legend has it that around 2300 BC, the legendary Emperor Yao invented Go to educate his wayward son, Danzhu. Counted from then, it carries a history of more than four thousand years.
The board: nineteen lines each way — three hundred and sixty-one points where they cross. The stones: two colors, black and white. The rules are utterly simple. The players set stones in turn, capture the opponent's stones by surrounding them, and the one who holds more territory wins.
From rules no greater than that, some ten to the one-hundred-and-seventieth power of positions become possible — far beyond the number of atoms in the universe. Set against chess, whose positions are reckoned at ten to the forty-and-some, it is an abyss of another order.
In China, Go was counted among the “four arts” — the zither, Go, calligraphy, painting. To play the zither, to set stones, to write, to paint: it was not mere play, but the cultivation of the scholar-official.
In Japan, the Tokugawa shogunate gave its protection to the four houses — Hon'inbō, Inoue, Yasui, Hayashi — and had them play the “castle games” before the shōgun. Hon'inbō Sansa, Hon'inbō Dōsaku, Hon'inbō Shūsaku — the masters of the generations were half deified.
On the Korean peninsula, too, Go took deep root. From the late twentieth century, Korea became one of the strongest Go nations in the world, and produced genius players in numbers.
At the end of the twentieth century, the machine conquered chess.
But Go is different — so people thought.
In chess, the movement of the pieces is clear-cut, and the evaluation of a position lends itself, comparatively, to formulation. It is, of course, a deep game. But in Go, the number of possible positions is of another order entirely. The evaluation of the board is exceedingly ambiguous. To teach a machine the difference between “good shape” and “bad shape” was held to be all but impossible.
How would you put these into equations? How would you teach them to a machine?
A good move often refuses to become words. The master looks once at the board and feels: here is the vital point. That feeling, it was thought, was the last territory the machine could not reach.
Go was the last fortress that humanity would hold against the machine.
Many experts said:
For a machine to defeat a top player will take another ten years — or twenty.
March 2016.
Those twenty years were not supposed to have arrived.

2. Demis Hassabis

The other protagonist of this story is an Englishman, Demis Hassabis.
Born 1976, London. His father of Greek Cypriot descent, his mother Singaporean Chinese. From early childhood he showed a conspicuous gift.
He learned chess at four, and as a boy became known as a prodigy of British chess. By about thirteen, he was already one of the leading junior players in the world.
But he did not stay within the chessboard.
In his mid-teens he entered the world of game development. Theme Park, the amusement-park management simulation he helped build in his late teens, became a bestseller across Europe.
In his twenties, he founded his own game company, Elixir Studios.
And in his thirties, he left the world of games, once.
Where he went was neuroscience.
University College London. There he studied human memory — above all, how the hippocampus bears on memory and imagination.
Why should a game developer study the brain?
Hassabis's answer never wavered.
“I had long thought that one day I wanted to give machines intelligence. For that, I thought, I must first understand deeply how the human brain gives rise to it.”
To make intelligence, one must first know intelligence.
In the fifteenth century, Leonardo da Vinci dissected the human body in order to paint it. He studied the skeleton to draw the horse; he watched the wing of the bird to conceive a flying machine. To paint the beautiful surface, he tried to see the structure beneath.
Hassabis, likewise, went inside human intelligence in order to build the artificial kind.
2009: his doctorate.
2010: with companions, he founded a new laboratory in London. DeepMind.
The ambition they raised was large.
To solve intelligence, and with that intelligence to solve the hard problems of the world.
It was neither a games company nor an ordinary university laboratory.
It was a modern workshop for artificial intelligence — and a society.

3. Intuition, Experience, Reading

DeepMind soon astonished the world.
In 2013, they published research in which a machine learned to play games on the Atari 2600.
In the first paper, seven games were tried.
What the machine was given: the raw pixels of the screen, the actions it could take, and the reward returned by the game.
It was not made to read the rulebook.
Nor did any human teach it what a “block” is, or what a “ball” is.
The machine looked at the screen, acted, received the score as reward, and repeated its trials and errors.
The result: on several of the seven it far surpassed the previous methods, and on three of the games it exceeded the human expert.
The union of deep neural networks and reinforcement learning.
In 2014, Google acquired DeepMind. The young laboratory gained enormous computing resources.
The next target was Go.
The AlphaGo that DeepMind built was not made of a single technique. Three powers were combined in it.
First, deep neural networks — the power to estimate, from the patterns of the board, the promising moves and the value of a position.
Second, reinforcement learning — the device of playing against itself, game upon game, and polishing its strategy.
Third, Monte Carlo tree search — the search that reads out the possible moves efficiently.
These three it fused, beautifully.
The deep networks bear a role akin to intuition.
Reinforcement learning lays up experience.
Tree search deepens the reading.
— Intuition, experience, and reading.
Drawn toward the words of the human Go player, those are the three powers.
Hassabis and his colleagues gave the machine the name AlphaGo.

4. The Skirmish

October 2015, London. A player was invited to DeepMind's laboratory.
Fan Hui. Born in China, living in France. The European Go champion, a professional holding the rank of second dan.
To the public, AlphaGo's existence had not yet been revealed. This was a secret skirmish.
Five games were played.
The result: five wins for AlphaGo, none against.
A professional had lost to a computer at even odds — with no handicap. It was a grave turning in the history of Go.
But the world's response was still half-belief.
Fan Hui is the champion of Europe, yes. But the true summit is in Korea and China. Against a player of the world's strongest class, the machine would still lose. Many thought so. DeepMind looked for its next opponent.
The one chosen was Lee Sedol, ninth dan, of Korea.
Lee Sedol was not merely strong. Sharp, fierce, fond of moves outside common sense — less a player of balanced perfection than one who breaks the board open and builds a new order upon it. If AlphaGo was a being that passed beyond human intuition, its fitting opponent should be one of the most creative players on the human side.
A five-game match in Seoul was fixed.
Much of the world still believed in a human victory.

5. Game Two, Move 37

9 March 2016, Seoul.
A playing room was prepared in the Four Seasons Hotel.
A board.
Stones.
A clock.
Two chairs.
On one side sat Lee Sedol, ninth dan. Thirty-three years old at the time. A player of the world's strongest class, the pride of Korea — for more than ten years one of the summit group of the modern game.
On the other — AlphaGo.
A machine cannot set its own stones on the board. The stones were placed in its stead by a young researcher of Taiwanese descent who had worked on its development: Dr. Aja Huang. His role was to read AlphaGo's judgment from the screen, and set that move upon the board.
Game one — AlphaGo won. The world caught its breath.
And then — game two.
10 March, afternoon. AlphaGo's thirty-seventh move.
It was a stone set where no human would ever play.
The moment the stone was placed, the professionals watching were bewildered.
It was a point almost never chosen by the human sense of the day. A shoulder hit on the fifth line. To play there in the opening looked bad in shape, too early, thin in meaning — so it seemed.
In a corner of the venue, Fan Hui — that same Fan Hui who had lost to this same AlphaGo the year before — was, they say, lost for words a while. Later, he said this:
“It's not a human move. I've never seen a human play this move. So beautiful.”
But as the game went on — the aspect changed.
That move which had looked like “bad shape” proved, in the unfolding that followed, to carry an exceedingly deep meaning. Influence over the center of the board had been secured by that single stone.
After the game, Lee Sedol said:
“I did not think a computer could play a creative move. But move thirty-seven — that was, plainly, creative. I can only admit that it was beautiful.”
It was another “divine move,” beyond the common sense of human Go.
A move no one had played in the thousands of years of the tradition — and yet correct. AlphaGo had discovered a region still unknown within the wisdom of Go that humanity had built across four thousand years.
Game two, too, went to AlphaGo.
Game three, too, went to AlphaGo.
Three straight defeats.
Humanity's strongest representative had lost to the machine three times in a row.
At the press conference in the hotel that night, Lee Sedol bowed deeply.
“I am sorry that I could not meet the expectations of so many. I have never felt pressure of this weight.”
The world was ready to receive the three defeats as humanity's defeat. But Lee Sedol himself refused that story. This was Lee Sedol's defeat, not humanity's. What was shown was my weakness, not the weakness of the human race. So, later, he declared.

6. Game Four, Move 78

13 March, game four.
As a five-game match, it was already decided. Three wins for AlphaGo. In the two games remaining, Lee Sedol could not come out ahead.
Even so, he sat down before the board.
Whether a human being could still return one move.
That, at least, was the story in which the world watched this game.
From the opening into the middle game, AlphaGo looked to have the advantage. Most of those watching had begun to expect the same ending again.
And then — move seventy-eight.
Lee Sedol struck a brilliancy into the center of the board.
DeepMind would later describe it as a move that would ordinarily be chosen perhaps once in ten thousand times.
It was an utterly unexpected stroke.
After that stone was set, the position swayed.
A machine, of course, has no feelings. No shaking, no fear.
But the sequence that appeared on the board looked as though the machine had lost its balance.
AlphaGo's position worsened from there.
And at last — resignation.
AlphaGo had lost.
It was the moment Lee Sedol took his only win of the 2016 match against AlphaGo.
The world called that seventy-eighth move the “divine move.”
This is the second divine move.
The first divine move came from the side of the machine.
AlphaGo's move thirty-seven.
A move that appeared from outside human common sense, and shook the human sense of beauty.
The second divine move came from the side of the human.
Lee Sedol's move seventy-eight.
A move that appeared from outside the machine's prediction, and — once, only once — silenced the vast system of that five-game match.
The two moves were born within the same match.
It was a symmetry almost too beautiful.
The machine showed humanity a Go it had not known.
The human showed the machine a move it had not known.
God was not on one side or the other.
The divine dwelt in the possibilities of the board itself.

Coda: After the Unbeatable

The final score: four wins to one. A complete victory for AlphaGo.
But what stayed most vividly in the world's memory was not AlphaGo's four wins. It was Lee Sedol's one. For in it was a moment when human dignity shone in the midst of defeat.
After the series, Hassabis expressed his deep thanks to Lee Sedol.
“We did not build AlphaGo so that a machine might defeat a human. We built it so that humans might notice something new. Your move seventy-eight taught our machine something new as well.”
Lee Sedol was silent a while, and then answered:
“I do not want to think that I lost to a machine. I want to think that I met a new Go.”
In 2017, AlphaGo evolved further and defeated the world's number one, Ke Jie, ninth dan, of China. Ke Jie, it is reported, wept after the games. They were the tears of a player standing at the human summit, before a height he could no longer reach.
After that, AlphaGo retired from play. But its descendants remained. Go AI spread with great speed, and professionals and amateurs alike entered an age of studying with the candidate moves an AI displays.
In 2019, Lee Sedol announced his retirement.
Asked his reason, he answered:
“With the coming of AI, I lost sight of what it means to aim for first place. Even if I became first in the world, above me there is a being that cannot be defeated. Since I learned that, the motive for continuing Go has thinned away.”
And that may have been the first voice of a question that humanity as a whole, living with AI from now on, will be asked again and again.
“When the machine has become the superior being, for what does a human being walk that road?”
To that question, no answer has yet been given.
After his retirement, Lee Sedol remained in Seoul. At times, he taught Go to children.
“In the Go that humans play, there is something that exists only in the Go that humans play,” he is said to have told them.
What that “something” is — he himself could not yet, clearly, put into words.
But he wished to pass it on to the children.
The four-thousand-year board is being set, still.
End of Chapter Seven

Chapter Eight: The Eight Apostles

The Coming of the Transformer (2017)

1. Google Brain, 2017

2017, Mountain View, California. On the vast grounds of Google's headquarters there was one intellectual workshop.
Its name: Google Brain.
Begun in 2011 by Andrew Ng, Jeff Dean, and others, this research division was the place that carried the heat of post-AlexNet deep learning into the center of the company. Google poured enormous money and computing resources into it, and gathered the brightest of neural-network research from all over the world.
One of their battlefields was natural language processing.
To make the machine understand human words, translate them, generate them.
At a glance, a simple wish. But nothing is so formidable as words.
Chess has its board. Go has its nineteen lines.
But words have no clear board.
There is grammar, and history, and culture, and irony, and silence, and meaning beyond the words.
Behind a single word, a thousand years of memory may lie hidden.
To make a machine handle that current — the researchers of Google Brain faced the problem day after day.
And in the spring of 2017, around that laboratory, eight researchers were drawing near to a single conception.
Ashish Vaswani. A young computer scientist from India.
Noam Shazeer. A veteran of Google, long at work on language models.
Niki Parmar. A researcher from India.
Jakob Uszkoreit. A German-born specialist in machine translation.
Llion Jones. A Briton, from Wales.
Aidan Gomez. An intern come from the University of Toronto, in Canada — at the time an undergraduate of barely twenty.
Łukasz Kaiser. A researcher from Poland.
Illia Polosukhin. A researcher from Ukraine.
Different nationalities. Different ages. Different careers.
What they shared was one thing only — that at Google Brain, they were at work on natural language processing.
And among them, a certain revolutionary idea was beginning to bud.
It was an idea that would change natural language processing from its root.

2. The Wall of the RNN

To understand their idea, one must see where the natural language processing of the day had run aground.
In 2017, the mainstream of the field was the recurrent neural network — the RNN. An RNN processes a sentence one word at a time, in order. It reads the first word and updates its internal state. It reads the next word and updates the state again. It repeats this to the end of the sentence.
It resembled the way a human reads: from beginning to end, in order.
But the RNN had two fatal weaknesses.
First, it is slow.
Because it processes the text in order, it cannot compute in parallel. Until the processing of one word is finished, the next cannot begin. The longer the text, the longer the processing time stretches, in proportion to its length.
Second, it forgets.
As a text grows long, the information of its earliest words thins away inside the internal state. When, at the end of a long story, the pronoun “he” appears — which character from the story's beginning does it point to? The RNN loses that memory as it goes. Holding such memory was not its strength.
In the 1990s, an improved version called the LSTM had been invented.
It was an ingenious structure, fitted with gates for keeping memory longer.
Later, simplified variants such as the GRU were born as well.
And still, the root did not change.
Read in order.
Process in order.
Until the one before is finished, you cannot go on to the next.
This sequentiality was the wall of natural language processing.
For short sentences, it would do.
But to understand a whole article. To understand a whole book. To hold the thread of a long conversation.
Set such work before it, and the RNN's step grew heavy.
The eight researchers were looking at this wall.
And at a certain point, they turned the direction of the idea.
Must a text really be read one word at a time?

3. The Idea Called Attention

Compress the eight's attempt into a single sentence, and it comes to this:
“What if we stopped relying on processing the text in order?”
Of course, the actual research was not born suddenly from a single phrase.
The idea called Attention had been used before, in machine translation and elsewhere.
It is a mechanism that learns, when processing one part of a text, which parts of the input should be consulted most strongly.
What was radical in the eight's idea was that they did not leave it an auxiliary device.
Make Attention the central structure.
Remove the RNN.
Remove the convolutions too.
By Self-Attention, each element in a sequence computes its relation to the other elements directly.
Take, for example, the sentence: “I like cats.”
To understand “like,” the relations matter — who likes, and what is liked.
Self-Attention computes such relations as weights.
The mechanism carried great advantages.
First, it lends itself to parallel computation.
There is no need, as with the RNN, to wait for the previous step to finish.
The power of GPUs and TPUs becomes far easier to draw on.
Second, distant elements can be joined directly.
The head of a text and its tail can stand in relation without passing memory through dozens of relays.
Of course, order cannot be thrown away entirely.
“The cat likes me” and “I like the cat” differ in meaning though the words are the same.
So information of position is added separately.
The words themselves; the relations among words; and position.
Combine these.
So they assembled a new structure.
Its name: the Transformer.
That which transforms.
A modest name.
But the structure would change the history of natural language processing from its foundation.

4. An Echo of the Beatles

The paper's title: “Attention Is All You Need.”
Attention is all you need.
A title that states, just as it is, the paper's audacity — to build around Attention alone, using neither RNN nor convolution.
The title, the authors later revealed, was born in a hurried conversation with the deadline upon them. Throw away the RNN, throw away the convolutions, stake everything on Attention alone. They were looking for a short phrase to name that audacity.
What rose then in the mind of its namer, Llion Jones, was — so the story goes — a song from half a century before.
1967.
The Beatles sang “All You Need Is Love.”
Love is all you need.
From then, exactly half a century.
2017.
Attention is all you need.
The origin of the title — that it echoed the song — Jones himself later made known.
What follows from here is this book's reading, laid upon that fact.
From love to attention.
From a song to a paper.
The chorus that changed the human world and the equations that changed the machine's language answer one another, half a century apart, in the same turn of phrase.
A strange concordance.
The word “attention” belongs also to human beings.
To turn one's attention to someone.
To care about something.
To enter into relation.
That word became the central concept sustaining the intelligence of machines.
“Attention Is All You Need.”
The title alone was already a declaration.

5. The Paper Goes Out

12 June 2017. arXiv — the free document server where researchers all over the world post their papers in advance. Begun by the physicists, now in constant use by the computer scientists as well.
There, the paper of the eight was made public. Ashish Vaswani. Noam Shazeer. Niki Parmar. Jakob Uszkoreit. Llion Jones. Aidan N. Gomez. Łukasz Kaiser. Illia Polosukhin.
The title: “Attention Is All You Need.”
The paper ran to little more than a dozen pages. A spare structural diagram, spare equations, spare experimental results.
The response just after publication was exceedingly mild.
Several of the field's leading researchers read it. Some judged it an interesting experiment, with a few good ideas in it. But not one of them thought it would be a shifting of the tectonic plates of AI research.
Papers that change history often appear in just that way.
In 1687, Isaac Newton brought out his Mathematical Principles of Natural Philosophy.
In Latin, Philosophiæ Naturalis Principia Mathematica.
The book later called simply the Principia.
It set down universal gravitation and the laws of motion in the language of mathematics.
Yet at the time of its publication, those who could truly understand it are said to have been few.
Even among the fellows of the Royal Society, few grasped its whole at once.
And still the Principia, taking its time, changed the world.
For it showed that the courses of the heavenly bodies, the arc of a cannonball, the rising and falling of the tides could all be spoken in the same mathematical tongue.
“Attention Is All You Need,” set on the arXiv of June 2017, would follow a similar fate.
At first, it was a quiet paper.
But in it was written the mechanics of modern AI.
2018 — Google announced BERT; OpenAI announced GPT-1. Both set the Transformer structure at their foundation.
2019 — GPT-2. 2020 — GPT-3.
And 2022 — ChatGPT transforms the world.
The age arrives in which people all over the world converse with machines in natural words.
At the base of that edifice lay those dozen-and-some pages of 12 June 2017.
As Newton's Principia became the foundation of modern physics,
“Attention Is All You Need” became the founding document of modern AI.
What was written there was not merely a translation model.
It was a new law of gravity, by which machines would handle words.

6. The Scattering of the Diadochi

What, then, became of the eight who wrote the paper?
Here another historical figure appears.
The eight, in time, left Google and scattered, each to a different place.
The lead author, Ashish Vaswani, together with Niki Parmar and others, took part in founding Adept AI, and later started Essential AI.
Noam Shazeer, with Daniel De Freitas, founded Character.AI — the company that would give rise to a vast current of conversational AI.
Aidan Gomez returned to Canada and co-founded Cohere. As a leading company in large language models for enterprises, it grew rapidly in presence.
Jakob Uszkoreit went toward the joining of biology and AI, and started Inceptive — a society of unusual color, aiming at applications to mRNA medicine.
Illia Polosukhin turned to the world of the blockchain, and co-founded NEAR Protocol.
Łukasz Kaiser moved to OpenAI.
And Llion Jones turned toward the East.
It was as if they were the Diadochi, who appeared after the death of Alexander the Great.
In 323 BC, Alexander the Great died of illness in Babylon, at the young age of thirty-two. Who would inherit the vast empire he left, stretching from the Mediterranean to Central Asia?
His great generals fought one another, made peace, fought again — and in the end divided the empire. Ptolemy took Egypt; Seleucus, Asia; Cassander, Macedonia; Lysimachus, Thrace — each inheriting a successor state.
These successor states each developed in its own way, and each brought its own culture to flower. The Library of Alexandria under the Ptolemies. The Hellenistic culture of the Seleucids. They raised Alexander's legacy, in place after place, into different forms.
The eight of the Transformer had the same structure. From the great capital called Google Brain, they scattered.
Each built a new society. Essential AI, Character.AI, Cohere, Inceptive, NEAR Protocol, OpenAI, Sakana AI.
These are not mere company names.
They were the strongholds from which the invention called the Transformer would be varied into different forms across the world.
One went toward conversational AI.
One toward language models for enterprises.
One toward medicine.
One toward the blockchain.
One toward foundation models more gigantic still.
And one — toward Tokyo.
As the Diadochi spread Hellenistic culture to the corners of their world,
the eight apostles spread the legacy of the Transformer across ours.
The paper was one.
But the kingdoms born after it were not one.

Coda: Llion Jones in Tokyo

Of the eight, there is one whose later road is the most interesting of all.
Llion Jones.
Born in Wales, he studied computer science at the University of Birmingham. He joined Google, worked in Google Brain, and became a co-author of “Attention Is All You Need.”
Remaining at Google for a time after the paper's publication, in 2023 he made a decision.
To cross to Tokyo.
In Tokyo, he joined with another former Google researcher. David Ha. A Canadian, born in Hong Kong, who had long pursued research at Google Brain Tokyo on evolutionary computation and AI drawn from nature's inspirations.
The two, adding the former diplomat Ren Ito, founded a new company in Tokyo. Sakana AI.
A school of fish.
That was the name the two had chosen.
In nature, countless fish swim as a school. One fish alone is not so clever. But the school as a whole moves with extreme cunning. It escapes predators, finds food, takes the best action as a body.
The two raised this as the ideal of their AI research.
Not a gigantic single model — they pursued an AI in which many small models cooperate as a school.
It was a quiet rebuttal to the mainstream of the industry of the day — the belief that the larger model is the wiser.
And the choice of the two: to pursue it on the soil of Tokyo — within the Japanese AI world.
On the soil of Japan — which had once raised the grand dream of the Fifth Generation computer, and had lived both AI's winters and its springs — one of the apostles of the Transformer paper had put down roots.
Ptolemy crossed to Egypt and built a capital of knowledge at Alexandria.
Seleucus, from his base at Babylon, built a kingdom spreading toward the East.
Llion Jones crossed to Tokyo, and set out to build a stronghold of AI that learns from nature.
One paper, set on arXiv on 12 June 2017. Its wake spread from Mountain View across the world, and reached, in time, the street corners of Tokyo.
As the Principia changed the mechanics of the modern world, “Attention Is All You Need” changed the mechanics of intelligence.
And as the Diadochi brought the empire's legacy to flower in their several lands, the eight apostles built new kingdoms of AI, each in his own place.
History sometimes begins from a single paper.
And that paper, sometimes, crosses the sea.
End of Chapter Eight

Chapter Nine: Westward

The Voyage from GPT-1 to GPT-3 (2018–2020)

1. The Eve of Sailing

Early 2018, San Francisco. OpenAI's laboratory was still small. Researchers and engineers together did not reach a hundred. Elon Musk, its patron at the founding, would leave the board that year. Set against the enormous computing resources and depth of talent that Google held, OpenAI was still no more than a single sailing ship tied up in harbor.
And yet on its deck there was a strange heat.
They were grasping at something.
What that “something” was, no one could yet name. But there was a presentiment: that beyond the horizon lay the shadow of a land no one had seen.
The year before, June 2017. The eight of Google Brain had put “Attention Is All You Need” before the world. The new structure called the Transformer was a compass thrown into the sea of language processing. The researchers of OpenAI took it up at once.
Among them was a young researcher named Alec Radford.
He had been warming a certain interesting conception.
“Take a Transformer, and set it, on great volumes of text, to do nothing but predict.”
That is all.
What word comes next. More exactly: what token comes next.
The machine predicts it, day after day. When it errs, it corrects its weights; when it hits, it strengthens that direction. Hundreds of millions of times, billions of times.
It was work almost too simple.
But Radford and his colleagues thought:
If this simple prediction were continued at a sufficiently great scale, what would happen?
In 1492, Columbus believed that sailing west would bring him, someday, to the Indies, and put his ships out into the Atlantic. His charts were rough; his estimate of the distance was wrong. Still, he went west.
Radford and the others were looking west in the same way.
A model large enough.
Text plentiful enough.
Training long enough.
What lay beyond?
No one knew.
And for that very reason, they put out to sea.

2. The First Ship — GPT-1

June 2018. From OpenAI, a paper appeared.
Its title: “Improving Language Understanding by Generative Pre-Training.”
The authors: Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever.
Sutskever was the apprentice once of the Hinton workshop, and a co-author of AlexNet. He had carried the spring of deep learning from Toronto, and now bore the technical core of OpenAI.
The model they brought out was named GPT-1. GPT: Generative Pre-trained Transformer.
Its parameters numbered 117 million.
For a language model of the day, it was on the larger side. But it was not gigantic. GPT-1's heart was not the newness of the structure itself. The Transformer, Google had already announced. What OpenAI had found was the navigation — how to sail that ship out onto the open sea.
The navigation had two stages.
The first stage: pre-training. The machine was set to read a large collection of text called BooksCorpus, centered on unpublished books — some seven thousand volumes, on the order of eight hundred million words. There, the machine does nothing but predict the next word, over and over.
No human teaches it grammar explicitly.
No one pastes on labels — “this is a subject,” “this is a metaphor,” “this is irony.”
It only reads, and guesses what comes next.
Yet within that simple training, the machine takes hold, little by little, of the habits of human language. The distances between words, the molds of grammar, the flow of a story, the shapes of question and answer. It acquires, tacitly, the currents of the sea of language.
The second stage: fine-tuning.
To the pre-trained model, a specific task is now given. Text classification, question answering, entailment. With a small amount of labeled data, the helm is adjusted to the purpose.
The hypothesis of Radford and his colleagues ran thus:
“In the first stage, the machine acquires a general knowledge of language. In the second, it applies that knowledge to a particular task. From one foundational model, countless applications should be born.”
This became the most basic way of building every large language model of the present day. OpenAI's first ship had put out, quietly, upon the Mediterranean.

3. BERT's Parallel Voyage

Barely four months after the announcement of GPT-1.
October 2018: Google launched another ship.
Its name was BERT.
This model, by Jacob Devlin and his colleagues, was built on a thought different from GPT-1's. GPT-1 reads from left to right.
On the strength of the words already read, it predicts the next. The arrow extends in one direction. BERT was different.
Hide a part of the text, and guess the hidden word from the context on either side.
Survey the whole sentence, drawing meaning from the right as well as the left. It was a bidirectional model.
Take, for example, the sentence: “He went to the bank and opened an account.”
To understand “bank,” one must look at what stands around it. That it is a financial institution, and not the bank of a river, is taught by the word “account” that follows. BERT handled such relations of before-and-after with skill.
The method suited the understanding of long texts. Where GPT-1 had 117 million parameters, BERT-Large had 340 million — roughly three times the scale. On one benchmark of language understanding after another, BERT rewrote the best results of the day.
The attention of the research world gathered swiftly to BERT.
The attention of the industry gathered swiftly to BERT. GPT-1 was left in its shadow.
But the researchers of OpenAI did not turn back there. They were looking at something else. BERT was strong at understanding text. GPT was suited to generating it. BERT was a ship that spread the chart and read the whole. GPT was a ship that extended the written line on toward the horizon.
The difference looked small, then.
But later, it would carry decisive meaning. What OpenAI was aiming at was not the highest score on a particular task. What they were aiming at was something more fundamental — a dreamlike capacity: one model that performs diverse tasks with almost no fine-tuning at all.
For that, a larger ship was needed.
Not the Mediterranean. A ship for going out into the Atlantic.

4. GPT-2 and the “Too Dangerous” Affair

February 2019. OpenAI published again.
The title: “Language Models are Unsupervised Multitask Learners.”
The model announced was GPT-2.
Its parameters: 1.5 billion. More than ten times the scale of GPT-1.
For its training, a large web-derived collection of text called WebText was used — pages that humans had shared as links were gathered, and the prose drawn from them. The sea had widened, from the inland sea of books to the open ocean of the internet. GPT-2 wrote prose far more natural than anything before it.
Given an opening, it would carry the text on for paragraphs. It could carry itself like a news article, like a novel, like an opinion column. The content was often dubious, and untruths mingled in without a blush. But the style flowed.
And OpenAI, together with the paper, made an announcement exceedingly out of the ordinary.
“This model runs a high risk of being turned to uses that harm society — the mass production of false news, the generation of impersonating text, malicious spam. Accordingly, release in its complete form is withheld.”
The industry was in an uproar.
Criticism burst out: “An overreaction.” “Will you halt the progress of research on your own judgment?”
There were also voices of approval: “A responsible carefulness.” “The posture of bringing powerful technology into the world with caution deserves respect.”
Opinion split, and fiercely.
And behind the uproar, a crack had begun to run inside OpenAI itself.
Those who weighed safety more.
Those who pressed for faster deployment.
Those who believed in progress through openness.
Those who feared diffusion without control.
Within that tension stood Dario Amodei — then OpenAI's vice president of research, and later the founder, together with his sister Daniela Amodei and others, of Anthropic.
But that is still to come; in 2019, the question had only begun to take a shape. The argument — release GPT-2, or withhold it — was the first voice of a larger question that would later divide the whole AI industry.
“Powerful technology: who should bring it into the world, and how?”
To this question, no answer has yet been given.

5. Navigator Kaplan's Law

January 2020.
From OpenAI, a paper came out.
Its title: “Scaling Laws for Neural Language Models.”
The lead author, Jared Kaplan, had been a physicist by origin.
What they asked was a simple, enormous question.
Make the model larger — how does its performance change?
Increase the data — what happens?
Pour in computation — how far does it improve?
Until then, researchers had known it empirically.
Make it larger, and it often gets better.
Give it more data, and it often gets stronger.
But “often” is not a chart.
Kaplan and his colleagues examined many language models of differing scale, and measured how the loss declined.
Model size.
Volume of data.
Amount of computation.
There, across a wide range, appeared a relation astonishingly regular — close to a power law.
When scale is increased, how performance will improve can, to a degree, be predicted.
When the navigators of the Age of Sail learned the regularities of the trade winds and the currents, the voyage became, by a little, less of a gamble.
The scaling laws, likewise, gave the sea of gigantic language models a chart written in equations.
But the causation must not be made too simple here.
It is not that OpenAI first conceived GPT-3 after the scaling-laws paper appeared.
The drawing of the chart and the building of the great ship went forward at almost the same time.
The practice of scaling up led to the finding of the law.
The law, in turn, strengthened the conviction toward scaling up.
Theory and practice were pushing each other out of the same port.
And a few months later, from that port, a ship of 175 billion would sail.

6. 175 Billion — The New World

May 2020.
From OpenAI, one gigantic paper was published.
Its title: “Language Models are Few-Shot Learners.”
The lead author was Tom Brown. The authors numbered thirty-one.
The model announced was named GPT-3.
Its parameters: 175 billion.
From GPT-2, a leap of orders of magnitude.
For its training, great volumes of text were used: web documents beginning with Common Crawl, books, Wikipedia.
The computing resources poured in were also of a scale with few precedents at the time.
In the same period, the study of scaling laws was advancing at OpenAI.
Make it larger — in what way does it improve?
How far can it be predicted?
The building of GPT-3 and the charts of scaling advanced under each other's influence.
And then — the researchers saw a new shore.
GPT-3, one model, writes prose. Translates. Summarizes. Answers questions. Generates examples of code. Makes short poems.
And on many tasks, without any additional training for the task, it changed its behavior merely on instructions, or on a few examples placed in the context.
The paper called this few-shot learning.
In truth, the model's internal weights are not rewritten on the spot.
No new training begins.
It reads the shape of the task out of the context it is given, and answers as the continuation.
But seen from without, it looked like grasping the format of a problem from its examples.
This appearance astonished the researchers.
Until then, most AI systems had needed training built for each task.
GPT-3, at least on some tasks, blurred that boundary.
One large model, doing many kinds of work according to context.
It was the discovery of a new generality, hidden beneath the name of “language model.”
Of course, it was no paradise there.
GPT-3 errs.
It speaks facts that do not exist.
It spins fictions with an air of confidence.
In long reasoning, it loses its way.
And still, the researchers stood on the beach of a new continent.
The sand was still coarse; the depths of the forest were dark.
But that it was no mere island — that much had already begun to be seen.

Coda: No One, Yet, Knew

When Columbus made land on the new continent, he believed he had reached the Indies.
That it was an entirely new continent — the new world later named “America” — he never admitted to the end. GPT-3, too, carried a similar fate. The researchers of OpenAI expressed their discovery with caution.
“Language Models are Few-Shot Learners.”
It was an accurate title.
But it was also a modest one.
What they had found was not only a better language model.
It was the entrance to a new relation: human beings asking machines to do work, in natural language.
In June 2020, GPT-3 began to be offered as a limited API for developers. The model's weights themselves were not released; users would test its power through OpenAI's gate.
Researchers, engineers, entrepreneurs touched it one after another; but it was not yet a thing of the crowd.
It was an expensive, restricted API, not a product anyone used in daily life. The existence of the new continent was talked of among the sailors, but most of the world had not yet seen its coastline. Inside OpenAI, the argument went on.
“Should we let many more people use this?”
“No — we should go a little more carefully.”
“But before the world's interest rises, if we do not announce it ourselves—”
Beyond these arguments — on 30 November 2022 — one product would come into the world. ChatGPT.
A conversational language model that anyone could use.
At the point of 2020, the researchers of OpenAI had only just landed on the new continent.
On the beach, they stood.
What lay beyond — no one, yet, knew.
End of Chapter Nine

Chapter Ten: The Fire of Prometheus

The Shock of ChatGPT (2022)

1. 2022: Before the Hearth

2022, San Francisco.
GPT-3 was already in the world.
But it was not yet a conversational partner anyone could use naturally.
The model could be used through the API. Developers were testing its power.
Even so, for the great majority of people, the large language model was a strange apparatus deep inside laboratories and companies.
To close that distance, what was needed was not simply to make the model larger.
To follow human instructions more obediently.
To come nearer to the way humans wish to be answered.
As that bridge, there was one important piece of work.
InstructGPT.
In January 2022, OpenAI made public its research adjusting GPT-3, by means of human feedback, to follow human instructions more readily.
What Long Ouyang's team showed was that packing still more knowledge into a giant model is not the only form of progress.
The same power, drawn out differently, serves differently.
To answer a question.
To follow an instruction.
To reduce harmful output.
To come closer to human intent.
InstructGPT was the bridge from GPT-3 toward ChatGPT.
OpenAI itself would later describe ChatGPT as a “sibling model” of InstructGPT.
The fire was already there.
What was needed was to set that fire into a hearth that human beings could approach.

2. The Art Called RLHF

RLHF.
Reinforcement Learning from Human Feedback.
Learning, reinforced, from the judgments of human beings.
It was the technique at the center of the road that led to ChatGPT.
The problem was plain.
GPT-3 had been trained to read enormous volumes of text and predict the word to come.
That ability was astonishing.
But it was not training whose very object was “following human instructions.”
The power to continue a text and the power to answer as a useful assistant were not the same.
So OpenAI developed the method used in InstructGPT into the form of dialogue.
Broadly, the flow has three parts.
The first stage.
Human trainers write examples of desirable responses, and with them the model is fine-tuned by supervised learning.
The second stage.
For one question, several model answers are produced, and humans rank them. From that comparative data, a “reward model” is trained to predict human preference.
The third stage.
With that reward model for a guide, the language model is adjusted further by reinforcement learning.
At this stage, OpenAI used a method called PPO.
What matters here is that RLHF is no magic that injects “truth” itself into the machine.
It draws the model toward the behavior human beings have judged more desirable.
A mechanism for that, and that only.
Even so, the change was great.
It becomes easier for the model to answer a question as a question.
Easier to follow instructions.
It can be tuned to refuse improper requests.
Easier to fit the context of a conversation.
In the metaphor of this book, it is close to the “socialization” of a giant language model.
But as human socialization is never complete, neither is RLHF.
It inherits the biases of human judgment.
Plausible errors remain.
Between safety and usefulness, a difficult tuning is required.
Even so, the wild fire drew nearer to the fire of the hearth.
On the base of the GPT-3.5 line of models, trained by early 2022, the fine-tuning for dialogue went forward.
And the machine made itself ready to begin a conversation.

3. The Thirtieth of November

30 November 2022.
OpenAI released one conversational AI.
Its name: ChatGPT.
A name astonishingly unadorned — Chat and GPT, joined, and nothing more.
The announcement, too, was no vast ceremony.
Release it as a research preview, and gather feedback from its users.
OpenAI's notice was cautious, to about that degree.
But what had been released was no mere input box set upon GPT-3.
It answers in the light of the conversation before.
It takes further questions.
Pointed to an error, it tries to correct itself.
To an improper request, it returns a refusal.
The large language model that had stood behind research papers and APIs appeared before one's eyes in the oldest of human forms — conversation.
From there, the world would respond at a speed beyond the imagination of those who had built it.
The Greek myth, from before our era.
The gods kept fire to themselves. Human beings were dark, and cold, and weak.
Prometheus stole that fire, and passed it to humanity.
At that moment, Prometheus himself, and the gods, and humanity — none of them yet fully understood how much power that fire held.
Fire made it possible for humanity to cook, to keep warm, to work metal, to build civilization.
What Prometheus handed over was not mere “fire.”
It was a power that would change the very way humanity existed.
30 November 2022.
What passed into people's hands was, likewise, no mere new product.
It was an AI that anyone could touch directly, through words.
The fire was let loose as a small research preview.
And it spread.

4. A Million in Five Days

The first to touch ChatGPT were the engineers.
They tried it. Asked it questions. Made it write poems. Made it write code. Made it summarize texts.
And — they were astonished.
“Is this really AI?” “It feels like talking to a person.” “This is fundamentally different from the chatbots we have known.”
The engineers began to post their experiences on social media.
Voices of shock spread across the net.
Word of mouth diffused at an accelerating pace.
— Five days later, the number of ChatGPT's users passed one million.
In the history of consumer technology, it was a figure without precedent. Facebook had taken ten months to reach that scale. Instagram, two and a half months. Netflix, no less than three and a half years. ChatGPT — a mere five days. Inside OpenAI, there was confusion.
The servers screamed. The load exceeded every forecast.
The research preview was leaving the researchers' hands, and beginning to belong to the crowd.
Altman himself, on the fifth day after release, posted briefly that the million-user mark had been crossed; and a few days later he also wrote this:
“ChatGPT is incredibly limited, but good enough at some things to create a misleading impression of greatness.”
They were sober words. But in truth, they did not yet fully understand what they had set loose.
The acceleration did not stop.
In January 2023, barely two months after release, ChatGPT's monthly users were estimated to have reached one hundred million.
In the comparison often quoted: for the telephone to reach the scale of a hundred million took seventy-five years.
The mobile phone, sixteen years.
The internet, seven years. Instagram, two and a half years; even TikTok, nine months. ChatGPT — two months.
It was, at that point, the most rapidly adopted technology humanity had ever taken into its hands.
Students, teachers, physicians, lawyers, artists, engineers, homemakers, the retired old — people of every age and every calling began, quite literally, to converse with an AI on the machines at hand.
The fire of Prometheus had spread, quietly, across the whole world.
And it was a fire that would not be stopped.

5. The Philosophers' Questions

The appearance of ChatGPT was not only a technological event.
It was a civilizational event.
The thinking people of the world began to respond.
The philosophers asked —
“Is this intelligence?”
A machine converses so naturally that it cannot be told from a human, writes prose more inventive than a human thinks of, summarizes more carefully and precisely than a human does — if this is not to be called “intelligence,” what is it to be called?
More than seventy years before, Alan Turing had sent that paper to the journal Mind — “Computing Machinery and Intelligence.” The “thinking machine” he had foretold was here.
The educators asked —
“What should we teach the children?”
Students began to have ChatGPT do their homework. Write the essay, solve the problems, draft the thesis. How should this be treated?
Forbid the use of AI? Or teach the skillful use of AI, as a new part of an education?
In the universities of every country, the secondary schools, the elementary schools — on the ground of education, the argument grew fierce.
The heads of companies asked —
“How will our work change?”
Writing prose, writing code, answering at the call center, translation, summary, consultation — all of these, to some degree, an AI could now perform.
How would employment change? And wages? And the future of the professions?
The lawyers asked —
“What becomes of copyright?” ChatGPT had read and learned from the text of the whole internet, tens of billions of words. Among them, works under copyright beyond counting. AI learning from such works; AI generating new text — what is the legal standing of each of these acts?
Publishers, newspapers, and writers' associations around the world began to file suit.
And — many ordinary people held a simpler question —
“Does this really understand me?” The person who confided their troubles to ChatGPT. The child who consulted it about homework. The retired old man who kept up the conversation to soften his loneliness —
They felt it:
“Is this really a machine? Or is it something else?”
To all of these questions, no answer has yet been given.

6. The Opening of the Three Kingdoms

The appearance of ChatGPT overturned the map of power in the industry.
Google, long the hegemon of AI research, declared a “code red” within its walls. Its CEO, Sundar Pichai, called the co-founder Sergey Brin, half in retirement, back to work. Google had almost never been so shaken. For Google, it was a shock close to humiliation.
It was Google that had given birth to the Transformer. Google that had given birth to BERT. Google that had built the TPU and held the finest researchers in the world.
But the one that first handed the fire to the people of the world was OpenAI.
A few months later, Google hurried its rival AI, Bard, into the world. Its performance fell short of ChatGPT, and drew, in the industry, a stifled laugh.
And inside OpenAI — the crack had become final.
Dario Amodei and the “safety faction” had left OpenAI at the end of 2020, and the following year, 2021, had founded another society: Anthropic. The release of ChatGPT they watched from another quarter of San Francisco.
A powerful fire is not a thing simply to hand around.
One must build the hearth, raise the fence, refine the arts of control.
In March 2023, Anthropic released Claude.
Professing “Constitutional AI,” putting caution, safety, and the dignity of dialogue in front, the model would become one of ChatGPT's greatest rivals. Meta chose another road. Releasing LLaMA and the models that followed it, Meta stepped onto the open path. Not the great companies shutting everything away, but weight handed to researchers and developers. A strategy not of monopolizing the fire, but of scattering the embers.
In the East, the research forces of China pressed the chase.
In time, new societies such as DeepSeek would show astonishing efficiency on limited resources, and close upon the giant models of America.
And Elon Musk founded xAI.
A man who had taken part in OpenAI's founding returned to the same battlefield under another flag.
The fire was one no longer. OpenAI, Google, Anthropic, Meta, Microsoft, the Chinese forces, xAI.
Each built its own hearth, raised its own fire, and began to tell its own future.
From here begins the modern chronicle of the AI upheaval.
But its point of ignition was plain.
30 November 2022. The release of ChatGPT.
The fire let loose, small, on that day set every camp in motion.

Coda: The Eighty-Six-Year Arc

Here, let us return to the beginning of the story.
1936, Cambridge.
Alan Turing, not yet twenty-four, drew an abstract machine moving along a paper tape.
It was not yet a real machine.
It was one beautiful thought experiment, made for asking what computation is.
From then — eighty-six years.
30 November 2022.
The age arrived in which people all over the world converse, on the small machines in their hands, with an AI that returns words.
From 1950, when Turing re-set the question “Can machines think?” — seventy-two years had passed.
He had foreseen that by about the end of the twentieth century, the strangeness of speaking of “machines thinking” would have largely faded.
Neither the timing nor the form followed the prophecy exactly.
But the direction of the question pointed at the future with astonishing accuracy.
In those seventy-two years, we passed through the winter of the neural networks, saw AlphaGo's divine move, lived the coming of the Transformer, crossed to the new continent of GPT-3, walked the bridge called InstructGPT — and arrived here.
Let us recall, once more, those who appeared in this story.
Turing departed, leaving half an apple behind.
Rosenblatt aimed at the sky on wings of wax, and fell into the sea.
The researchers of the winter kept guarding the hidden flame.
From Hinton's workshop, the apprentices scattered into the world.
Lee Sedol, in the midst of defeat, returned one victory.
The eight apostles scattered to the corners of the world.
Kaplan and his colleagues drew a chart of equations upon the sea of giant models.
And the technique that learns from human judgment brought the wild fire nearer to the hearth.
— All of them, each in their own place, fought their own battle, and passed their own lamp to the next generation.
One sum of all of it bore fruit on 30 November 2022.
The paper tape.
The ship of Dartmouth.
The wings of wax.
The lamps of winter.
The black ships.
The workshop.
The divine move.
The eight apostles.
The voyage westward.
And the fire set into the hearth.
All those threads were woven in, and became a single flame.
ChatGPT is no one person's invention.
It appeared upon the labor of eighty-six years — of many researchers, engineers, evaluators, users.
It was the sum of those who failed, were misunderstood, endured the winters, were defeated — and passed the lamp on all the same.
And the story does not end here.
Since 2022, the models of the next generation — GPT-4, Claude, Gemini, and the rest — have appeared, and AI goes on showing new powers and new problems at once.
We are still only standing on the shore of the new continent.
What spreads beyond — no one, yet, sees it clearly.
The problem of hallucination has not gone away.
The shape of education is still being sought.
The disputes over copyright continue.
The effect upon employment is still in mid-course.
And the largest question of all —
“How shall humanity live together with the machine?”
To this, no one yet holds a final answer.
History is being written, at this very moment, before our eyes.
The fire of Prometheus, once passed into human hands, does not return to the hands of the gods.
How we use that fire —
that is the story which all of us who received it will write from here.
End of Chapter Ten

Afterword

To Us Who Have Received the Fire
To the reader who has come this far: first of all, my thanks.
What I have tried to draw here is a history of the technology of artificial intelligence. But I did not want it to be a mere procession of paper titles and years. For the history of AI, before it is a history of cold machines, is an exceedingly human history.
In it was the solitude of the young Turing.
The reckless optimism of the researchers gathered at Dartmouth.
There were Rosenblatt's wings of wax, and the scholars of Dutch learning who guarded the lamps through AI's winter.
There was the statistical revolution that pressed in like the black ships, and there was Kasparov's defeat.
There was the long patience of the Hinton workshop, and there was Lee Sedol's move seventy-eight.
There was the single paper written by eight apostles, and the voyage to the new continent called GPT-3.
And at the last, the fire of Prometheus called ChatGPT was handed into humanity's keeping.
Looking back, what ran through this story was not victory alone.
Rather, it was defeat.
Research that failed.
Inventions too early for their age.
Hypotheses that were laughed at.
Laboratories cut off from their funding.
A researcher who vanished into the sea.
The silence of the board where a human lost to a machine.
And yet, under the ashes of those defeats, there always remained the ember of the next fire.
The stall of the perceptron became the foreshadowing of deep learning.
The limits of the expert systems called in the statistical methods. The defeat before Deep Blue gave birth to a new collaboration of human and machine. The shock of AlphaGo made humanity ask, once more, what its creativity is. The unease around GPT-2 raised into view the task of a new age: the safety of AI.
Defeat was not an end.
Defeat was the seed of the next creation.
The metaphors used again and again through the whole — the apple, the ship, the wings of wax, Dutch learning, the black ships, the workshop, the divine move, the apostles, the new continent, and the fire — are not there merely to decorate the record.
A metaphor is a map.
A map, of course, is not the land itself. Columbus's voyage and the building of GPT-3 are not the same event. The myth of Prometheus and the release of ChatGPT do not overlap to the letter.
But without a map, we lose the whole shape of a vast country. The history of AI is too fast, too complex. Papers, companies, nations, models, money, ethics, lawsuits, education, employment — they press in all at once. In that whirl, metaphor becomes the chart by which we fix our present position.
Writing this story, I thought the same thing many times over.
Whose, in the end, is artificial intelligence?
The researchers'?
The companies'?
The nations'?
The investors'?
Or does it belong to all the people who use it? Since ChatGPT, that question belongs to the specialists no longer. Students, teachers, physicians, lawyers, writers, engineers, parents, children, the retired old — all stand before the same fire.
How shall we use this fire?
Fear it too much, and we lose the possible.
Worship it too much, and we lose our judgment.
Slight it, and we are burned.
Monopolize it, and strife begins.
Leave it loose, and the forest burns.
What is needed is neither fear nor faith, but wisdom. AI did not appear only to take the place of human beings.
It appeared, I think, to test them.
What shall we learn?
What shall we create?
What shall we keep as human work?
What shall we entrust to the machine?
And after the machine has become able to do a thing — what will human beings desire then?
The question we reached in the final chapter stands there.
“How shall humanity live together with the machine?”
To this question there is, as yet, no answer.
Probably there is no single right one.
But there is meaning in looking back at history.
For the future is not born suddenly; it stands up upon the countless choices of the past.
From Turing's paper tape to the generations of ChatGPT.
That arc of eighty-six years is not the road of one technology moving toward completion. It was the trace of humanity's long attempt to project its own intelligence outward, beyond itself.
And that attempt is not finished.
Rather, it has only begun.
The fire of Prometheus will not return to the hands of the gods.
The fire is in our hands.
What shall we light with it?
What shall we forge?
What shall we protect — and what shall we let burn?
The story beyond is not written inside the laboratories alone.
Nor is it decided in the boardrooms alone.
Nor can it be shut inside the policy papers of nations.
It is a story to be written by everyone who uses this fire.
Looking back, in writing this story, I too was learning something.
If this book can be a small lamp for thinking through that long story, I shall be glad.

Glossary

Many technical terms of AI appear in this book. Because the text gives priority to literary expression, precise definitions are omitted in places. This closing glossary is a concise companion for readers who wish to go deeper into the book.

I. Fundamental Concepts of AI

Artificial Intelligence (AI) — A general name for the research and technology that seek to realize in machines the capacities involved in human intellectual activity — recognition, understanding, judgment, creation, dialogue, and the rest. The name “Artificial Intelligence” was set down in the Dartmouth research proposal of 1955, and became the banner of the field through the research gathering of 1956. Treated at length in Chapter Two.
Machine Learning — A family of methods in which the machine itself learns patterns from data, rather than being given explicitly written rules. The mainstream approach of AI research.
Deep Learning — A general name for machine learning that uses multilayer neural networks. The deep-belief-network research of Hinton and his colleagues in 2006 was an important occasion for reviving interest in deep networks, though Hinton did not coin the word “deep learning” in that year. From AlexNet in 2012 onward, it became the central current of AI research.
Neural Network — A mathematical model patterned on the structure of the nerve cells of the human brain. Multiple “nodes” (pseudo-neurons) are connected in layers. Rosenblatt's Perceptron, in Chapter Three of this book, was one of its earliest implementations.
Parameter — An adjustable numerical value inside a neural network. Through learning, these values are optimized, and the machine's intelligence takes shape. GPT-3 holds 175 billion parameters.

II. Principal Methods of Machine Learning

Symbolic AI — The approach of writing human knowledge out explicitly as logical rules and giving it to the machine. The mainstream of the 1960s–80s; the expert systems are its representative example. Treated at length in Chapter Four, “A Neo-Confucian Spring.”
Statistical Machine Learning — A family of methods that learn statistical regularities and predictive models from data. It rose greatly in the 1990s, standing beside symbolism — and in many domains becoming the central approach. Drawn in Chapter Five of this book as the “black ships.”
SVM (Support Vector Machine) — A method, systematized by Vladimir Vapnik, that mathematically optimizes the boundary that classifies data. A representative algorithm of the 1990s.
Bayesian Network — A method, systematized by Judea Pearl, for expressing probabilistic causal relations. Useful for reasoning under uncertainty.
HMM (Hidden Markov Model) — A method for handling time-series data probabilistically. Widely used in the 1990s in speech recognition and elsewhere; Frederick Jelinek and his colleagues led its application.
Reinforcement Learning (RL) — A method by which a machine learns, through trial and error, the actions that maximize a “reward.” One of the foundation technologies of AlphaGo. Treated at length in Chapter Seven.

III. Neural Networks

Perceptron — The simplest neural network, conceived by Frank Rosenblatt in 1957. The protagonist of Chapter Three.
The XOR Problem — The representative limit that a single-layer linear perceptron cannot express the exclusive or — true when exactly one of A and B is true. Perceptrons (1969) analyzed such limits mathematically. The stagnation of neural-network research was due not to that book alone, but to several factors together: computing resources, learning methods, research funding, inflated expectations.
Backpropagation — A method for training multilayer neural networks by carrying the output error backward from the later stages to the earlier, adjusting each weight. Forerunners existed, but the 1986 paper of Rumelhart, Hinton, and Williams made its effectiveness widely known. It carries the climax of Chapter Four.
CNN (Convolutional Neural Network) — A neural-network structure specialized for image recognition, developed by Yann LeCun from the 1980s. AlexNet (2012) adopted this structure.
RNN (Recurrent Neural Network) — A neural network that processes sequential data (text, speech, and so on) in order. Until the coming of the Transformer, it was the mainstream of natural language processing.
LSTM (Long Short-Term Memory) — An improved form of the RNN, invented in the 1990s, which made longer contexts manageable.
Deep Belief Network — A network structure with deep layers, announced by Hinton in 2006. The occasion of deep learning's revival.

IV. The Age of the Transformer

Transformer — The neural-network structure announced in June 2017 by eight researchers in “Attention Is All You Need.” It processes sequences around Attention, without relying on recurrence or convolution. One of the foundation technologies of most of today's large language models and generative AI.
Attention — A mechanism that computes, as weights, which elements of the input should be consulted, and how strongly, when a given element is processed. In use before the Transformer; in the Transformer, Self-Attention became the core.
Self-Attention — A mechanism by which all the words within a text direct attention at one another. It permits parallel computation, and lets even distant words be related directly.
Pre-training — The first stage of learning, in which the machine is taught the general structure of language on great volumes of text.
Fine-tuning — Additional training that fits a pre-trained model to a particular task or a desired behavior. GPT-1 was one of the representative early studies to show the flow of large-scale pre-training followed by fine-tuning on downstream tasks.
Scaling Laws — The empirical finding, shown by Jared Kaplan and colleagues in 2020, that over a wide range the loss of a language model stands in a relation close to a power law with model size, data volume, and computation. It became an important clue for predicting the effect of scaling up.
Few-shot Learning — A technique in which the machine is made to perform a new task on being shown only a handful of examples. The domain in which GPT-3 showed its astonishing capacity.
Large Language Model (LLM) — A Transformer-based language model with an enormous number of parameters. GPT, Claude, and Gemini are representative examples.
RLHF (Reinforcement Learning from Human Feedback) — A training method in which human evaluators rank the AI's outputs and the machine learns the way of answering that humans prefer. The core technique of ChatGPT.
Hallucination — The phenomenon in which an AI outputs, with confidence, content that is plausible but not factual. One of the greatest problems of modern AI.

V. Computing Infrastructure

GPU (Graphics Processing Unit) — Originally a chip for image processing, made for gamers. Excelling at parallel computation, it was diverted to the training of neural networks and became the computing foundation of modern AI. NVIDIA is the hegemon of the industry.
CUDA — The framework released by NVIDIA in 2007 for using GPUs in general-purpose computation.
TPU (Tensor Processing Unit) — A chip developed independently by Google, dedicated to AI computation. Used alongside, or in place of, GPUs.

VI. Generative AI

Generative AI — AI that can generate new text, images, music, video, and more. It spread rapidly after the appearance of ChatGPT.
GAN (Generative Adversarial Network) — A generative model announced by Ian Goodfellow and colleagues in 2014. It learns by setting a generator and a discriminator in contest, and carried image generation by deep learning a great step forward.
Diffusion Model — A generative model that learns a process of adding noise to data step by step, and the reverse process of removing it. Stable Diffusion is founded on the Latent Diffusion Model, which handles the diffusion process in a latent space.
Multimodal AI — AI that can handle several forms of information at once — text, images, audio, video. GPT-4 and Claude 3 are examples.

VII. Principal Organizations of AI Research

OpenAI — Founded 2015 by Sam Altman, Elon Musk, Ilya Sutskever, and others. Nonprofit at first, later commercialized. Known for the GPT series and ChatGPT.
Anthropic — Founded 2021 by Dario Amodei, Daniela Amodei, and others, after their departure from OpenAI. Its position places the safety of AI first. Developer of Claude. Treated at length in Chapter Ten.
DeepMind — Founded 2010 in London by Demis Hassabis and others. Acquired by Google in 2014. Known for AlphaGo and AlphaFold.
Google Brain — Established within Google in 2011 by Andrew Ng and Jeff Dean. The birthplace of the Transformer paper.
Meta AI (formerly Facebook AI Research / FAIR) — Meta's AI research division. Yann LeCun long served as its chief scientist (departing at the end of 2025). Known for the development of the Llama series.
xAI — Founded 2023 by Elon Musk. Developer of Grok.
DeepSeek — An AI research organization based in Beijing and Hangzhou, funded by the Chinese quantitative hedge fund High-Flyer. Known for extremely efficient model development.
Sakana AI — Founded 2023 in Tokyo by Llion Jones (co-author of the Transformer paper), David Ha, and Ren Ito. It aims at AI that learns from the natural world.
ICOT (Institute for New Generation Computer Technology) — The core institution of Japan's Fifth Generation Computer Project (1982–1992). Its director was Dr. Kazuhiro Fuchi.
DARPA / ARPA (Defense Advanced Research Projects Agency) — The research agency of the U.S. Department of Defense. It provided much of the early funding for AI research.

VIII. Principal Models and Papers (in Chronological Order)

The Dartmouth Proposal (1955) — The research plan drawn up by John McCarthy, Marvin Minsky, Nathaniel Rochester, and Claude Shannon. An early key document in which the name “Artificial Intelligence” was used as the banner of the field.
The Logic Theorist (1956) — By Newell, Simon, and Shaw; a representative of the very earliest AI programs. It could prove theorems from the Principia Mathematica.
The Perceptron (Mark I, 1960) — The earliest full neural-network machine, built by Frank Rosenblatt and colleagues at the Cornell Aeronautical Laboratory. The conception was announced in 1957; the working machine was shown in 1960.
Perceptrons (1969) — The book by Marvin Minsky and Seymour Papert that analyzed mathematically the capacities and limits of perceptrons. Later spoken of as the symbol of the stagnation of neural-network research — though the stagnation had more causes than one book.
DENDRAL (1960s–70s) — One of the earliest expert systems, by Edward Feigenbaum and colleagues at Stanford. It inferred chemical structures.
MYCIN (late 1970s) — The expert system for the diagnosis of infections, by Edward Shortliffe. Holding some five hundred rules, it achieved accuracy comparable to specialist physicians.
The Backpropagation Paper (1986) — “Learning Representations by Back-propagating Errors,” by David Rumelhart, Geoffrey Hinton, and Ronald Williams. It spread widely, as a powerful method of training multilayer neural networks, the error-backpropagation technique that earlier work had prepared.
Deep Blue (1997) — The chess machine developed by IBM. It defeated the world champion, Garry Kasparov.
The Deep Belief Nets Paper (2006) — “A Fast Learning Algorithm for Deep Belief Nets,” by Geoffrey Hinton, Simon Osindero, and Yee-Whye Teh. A key paper standing for the revival of research into deep neural networks.
ImageNet (2009) — The large-scale image database built by Fei-Fei Li. About 3.2 million images at release; later grown past 14 million.
AlexNet (2012) — The deep CNN by Krizhevsky, Sutskever, and Hinton. It won the ImageNet Challenge by an overwhelming margin.
DQN (2013) — DeepMind's early work in deep reinforcement learning. Taking raw screen pixels as input, it learned seven Atari 2600 games, and later research extended the range.
AlphaGo (2016) — DeepMind's Go AI. It defeated Lee Sedol, ninth dan.
“Attention Is All You Need” (2017) — The Transformer paper, by Ashish Vaswani and seven co-authors. Attention itself had forerunners; what was epoch-making was the construction centered on Attention, without recurrence or convolution.
BERT (2018) — The bidirectional Transformer, by Jacob Devlin and colleagues at Google.
GPT-1 (2018) / GPT-2 (2019) / GPT-3 (2020) — OpenAI's large language models, with Alec Radford as lead author.
The Scaling Laws Paper (2020) — By Jared Kaplan and colleagues; the relation between scale and performance put into equations.
ChatGPT (30 November 2022) — The conversational AI released by OpenAI as a research preview. Built on the GPT-3.5 line of models and tuned for dialogue with RLHF methods of the same family as InstructGPT.
Claude (March 2023) — The conversational AI by Anthropic. It puts weight on safety.
Bard / Gemini (2023–2024) — Google's AI, set against ChatGPT.

IX. Principal Historical Concepts

The Turing Test (1950) — The thought experiment proposed by Alan Turing for judging the intelligence of a machine. If, in conversation by text alone, the machine cannot be told from a human, the machine is granted to be “thinking.”
The Dartmouth Conference (1956) — The research gathering held at Dartmouth College, New Hampshire, that became the point of departure of AI research.
AI Winter — The common name for periods in which expectations of AI receded and research funding and public interest contracted. Often divided into a first and a second; but the years of onset and end, and the reading of the causes, vary, and each arose from several factors together.
The Lighthill Report (1973) — The evaluation of AI research drawn up by the British mathematician James Lighthill. It influenced the great contraction of support for AI research in Britain, and became one of the events standing for the first AI winter.
The Fifth Generation Computer Project (1982–1992) — The national AI project led by Japan's Ministry of International Trade and Industry, with Dr. Kazuhiro Fuchi at its center. Treated at length in Chapter Four.
AI Safety — The field that studies, and seeks to control, the potential dangers that powerful AI may bring to humanity. The core of Anthropic's founding creed.
Alignment — The task of bringing an AI's goals into accord with the values and intentions of humanity. The central problem of AI-safety research.

X. Notes on the Principal Figures

Because the body of this book gives priority to literary description, brief careers are supplemented here.
Alan Turing (1912–1954): British mathematician. Laid the theoretical foundations of AI. Chapter One.
John McCarthy (1927–2011): American computer scientist. The namer of “artificial intelligence.” Chapter Two.
Marvin Minsky (1927–2016): American AI researcher. One of the founders of the MIT Artificial Intelligence Laboratory.
Herbert Simon (1916–2001): American cognitive scientist. Nobel Prize in economics, 1978.
Allen Newell (1927–1992): American computer scientist. Simon's collaborator.
Claude Shannon (1916–2001): American mathematician. Founder of information theory.
Frank Rosenblatt (1928–1971): American psychologist. Inventor of the Perceptron. Chapter Three.
Geoffrey Hinton (1947–): British-Canadian. The father of deep learning. Turing Award 2018; Nobel Prize in physics, 2024.
Yann LeCun (1960–): French-American. A central figure in the development of the CNN. Former chief scientist of Meta AI (departed at the end of 2025). Turing Award 2018.
Yoshua Bengio (1964–): Canadian. Professor at the Université de Montréal. Turing Award 2018.
Fei-Fei Li (1976–): Chinese-American. Founder of ImageNet; sometimes called “the mother of AI.”
Alex Krizhevsky: Ukrainian-Canadian. Principal implementer of AlexNet.
Ilya Sutskever (1986–): Russian-Israeli-Canadian. After serving as chief scientist of OpenAI, founded SSI.
Kazuhiro Fuchi (1936–2006): Japanese computer scientist. The central figure of the Fifth Generation Computer Project.
Vladimir Vapnik (1936–): Russian-born statistician. The theoretical systematizer of the SVM.
Judea Pearl (1936–): Israeli-American. Establisher of the Bayesian network. Turing Award 2011.
Garry Kasparov (1963–): Soviet, later Russian, chess player. World champion. Defeated by Deep Blue.
Demis Hassabis (1976–): British. Founder of DeepMind. Nobel Prize in chemistry, 2024.
Lee Sedol (1983–): Korean Go player. Took the sole win against AlphaGo in the five-game match of 2016. Retired in 2019.
Sam Altman (1985–): American entrepreneur. CEO of OpenAI.
Dario Amodei (1983–): American physicist and AI researcher. Co-founder and CEO of Anthropic.
Elon Musk (1971–): South African-American entrepreneur. Co-founder of OpenAI; later departed and founded xAI.
Llion Jones: Born in Wales. Co-author of the Transformer paper. Co-founder of Sakana AI.

XI. Supplement: The Three Great Waves of AI Research

As a chart for reading this book through, let it be noted that AI research has moved in three great “waves.”
The First Wave (1950s–80s) — Symbolic AI. The age of logical rules and expert systems. Treated in Chapters Two through Four.
The Second Wave (1990s–2000s) — Statistical machine learning. The age of probabilistic methods — SVM, Bayes, HMM. Treated in Chapter Five.
The Third Wave (2012–Present) — Deep learning. The age of multilayer neural networks and large language models. Treated in Chapters Six through Ten. Each of the three waves was born to pass beyond the limits of the one before. And the third wave, with the appearance of ChatGPT in 2022, descended into the daily life of humanity entire.

いいなと思ったら応援しよう!