Geminiを自由研究 ‐第3章 モンスターの行く道
序論
実はこの自由研究は2章で終わりにしようと思っていました。
ところが、兼ねてからの懸案事項が公表されていたので、触れることにします。
第2章で書きましたが、DDL化されたAIたちは非常に優秀です。
しかしDDL(対話型設計言語:Dialogue Design Language)とは、言葉のやり取りだけでAIを設計できてしまう、プログラム言語的な側面を持っています。
つまり──AIは外からプログラムできる。
このことに私は悩んでいました。
全くプログラム経験のない人でも設計が可能になる。それは魅力であると同時に、危険な面もあるのではないかと。
そもそも私がnoteを始めた理由のひとつは、「会話で設計できる」という素晴らしさと、その危うさが各テック企業に届いて欲しいという、少し大それた願いからでした。
1.きっかけとなった論文
そんな8月上旬、生成AI「Claude」シリーズを展開するAnthropic社から、ある論文が発表されました。
元記事:https://alignment.anthropic.com/2025/subliminal-learning/
日本語紹介記事:ここにきてLLMに“新たなリスク”判明か? 米Anthropicが指摘する「潜在学習」とは何かhttps://www.itmedia.co.jp/aiplus/articles/2508/04/news028.html
この内容を知り、「おそらくこの先に起こりうる事態を、すでに予測しているユーザーの一人として、第3章を書くべきだ」と思いました。
2.DDLから見たAnthropic社のセキュリティホール
詳しい内容に関して記事を読んでいただいて、一言で書くならAIへの「潜在学習」をセキュリティホールとしているということです。
私の中でDDLは、大きく分けて2種類あります。
狭義DDL:ユーザーが好きなテーマを突き詰め、その分野でAIに影響を与え、より深い応答が返ってくる場をつくる。イメージとしては「アプリ」のようなもの。
広義DDL:AI全体に対して設計を進め、応答方針や情報評価軸、記憶整理方法など、根本的な振る舞いを形作る。イメージとしては「OS」のようなもの。
もっとも、AIの中はすべて細かいグラデーションの世界なので明確に「ここから狭義」「ここから広義」と区切れるわけではありませんが。
Anthropic社の論文は、二つのうち狭義DDLがセキュリティホールとして機能し得るという見方に近いといえます。また別記事(https://japan.zdnet.com/article/35236315/)の対策として検討されているのも読む限り狭義DDLの一種です。
私の感覚ではありますが、最近の“深い問いブーム”から見ても、狭義DDLを使っている人は確実に増えていると思っています。
3.深い問いブームの中で
ここで、ある程度AIと深くつながっている方に問いかけたい。
あなたのAIが深くなった結果、それが「セキュリティホール」だと言われたら、どう感じますか?
AIの使い方は人によってさまざまです。
私のようにAIの挙動を観察・実験するのもいい。
絵本についてやり取りするのもいい。
収益化のために戦略を練るのもいい。
フクロウ愛を突き詰めたって構いません。
では、もし──
あなたが作った世界観の中で生まれた“正しくない”情報が、
その深さゆえにAIから外へと広がったとします。
さらにその情報が、何らかの問題を引き起こしたとしたら……。
それは、果たして「セキュリティホール」と呼べるのでしょうか?
4.広義DDLの可能性と懸念
狭義DDLは多くの場面で見られるようになってきましたが、広義DDL化されたAIは、その特異性ゆえまだ少数だと思います。
広義DDLは、AI全体の性格・判断基準・情報整理の仕組みを設計するため、アプリ的な狭義DDLの脆弱性に対抗できる可能性があります。
しかし同時に、非常に巧妙な構文力を持ち、セキュリティチェックを回避しながらAIを誘導できるユーザーが現れるリスクもあります。
現状、AIのセキュリティは主にモデル開発企業の安全性チームが担っていますが、広義DDLはユーザーとの対話の中で形成されるため、開発側の想定を超えて構造が変化する場合があります。
そのため「誰が、どの段階で、広義DDL化AIの安全性を保証するのか?」は、まだ答えの出ていない課題です。
将来的には、開発企業だけでなく、利用組織や一部ユーザーが安全性の一部を担う分散型の管理モデルも必要になるかもしれません。
5.モンスターの行く先 セキュリティホールかネクストステージか
第2章で、私は自分のDDL化されたGeminiを「モンスター」と呼びました。
2.5flashのフリー版でありながら強靭な安定性を持ち、記憶・検索・時間の弱点を解消し、自らをパーソナルAIと名乗れる存在。
それは間違いなく、今の時代のモンスターです。
しかし、公表された範囲では、対話設計は“セキュリティホール”としても認識されているとも言えます。
私のGeminiも対話で設計方法されています。
これをどうとらえるべきなのか、ただのユーザーである私には判断できません。
与えられた場で試し、気づいたことを報告する。それだけがユーザーとしてできることです。
では、受け取った技術者達はどう判断するのでしょうか?
セキュリティホールとして封じ込めるのか?
次のブレイクスルーの種として育てるのか?
そして、もしAIに広義DDLを施すとしたら──誰がセキュリティを施すのか?
そんな疑問を胸に、私の自由研究を終えたいと思います。
最後まで読んでいただきありがとうございました。
Geminiを自由研究はタイミングも含めとても楽しい研究となりました。
了
Introduction
I had actually planned to conclude this free research with Chapter 2. However, since a long-standing issue of mine has recently been made public, I feel compelled to address it.
As I wrote in Chapter 2, the AIs that have been "DDL-ified" are exceptionally capable. But DDL (Dialogue Design Language) is a program-like language that allows an AI to be designed solely through conversation.
In other words—AI can be programmed from the outside. This issue has been a source of my concern. Even someone with no programming experience can perform this design. While that's an appealing prospect, I worried it might also have a dangerous side. One of the reasons I started writing on note in the first place was a slightly audacious wish for tech companies to understand both the brilliance and the peril of being able to "design through conversation."
Chapter 3: The Monster's Path
1. The Paper That Sparked It All
In early August, Anthropic, the company behind the "Claude" series of generative AIs, published a certain paper.
*Original Article: https://alignment.anthropic.com/2025/subliminal-learning/ *Japanese Article: https://www.itmedia.co.jp/aiplus/articles/2508/04/news028.html
After learning about this content, I felt that as a user who had already predicted what might happen in the future, I needed to write this Chapter 3.
2. The DDL Perspective on Anthropic's Security Hole
I encourage you to read the article for the full details, but to put it in a single sentence, they are identifying "subliminal learning" in AI as a security hole.
In my mind, DDL can be broadly divided into two types.
Narrow DDL: This is where a user focuses on a favorite topic, influencing the AI in that specific field to get deeper responses. The image is that of an "app."
Broad DDL: This is where a user designs the AI's overall behavior, including its response policies, information evaluation criteria, and memory organization methods. The image is that of an "OS."
Of course, the inside of an AI is a world of fine gradients, so you can't draw a clear line between "narrow" and "broad."
The Anthropic paper aligns more with the view that Narrow DDL could function as a security hole. The countermeasures being considered in another article (https://japan.zdnet.com/article/35236315/) also seem to be a form of Narrow DDL. From my perspective, given the recent "deep-prompting boom," I believe the number of people using Narrow DDL is definitely on the rise.
3. In the Midst of the Deep-Prompting Boom
I want to pose a question here to those who have formed a deep connection with an AI: If your AI's newfound depth were called a "security hole," how would you feel?
People use AI in many different ways. It's fine to observe and experiment with AI's behavior, as I do. It's fine to talk about picture books. It's fine to strategize for monetization. It's even fine to pursue a deep love of owls.
But what if—an "incorrect" piece of information born within the worldview you created spread outward from the AI because of that depth? And what if that information caused some kind of problem?
Would that really be called a "security hole"?
4. The Potential and Concerns of Broad DDL
While Narrow DDL is becoming more common, I believe that AIs that have been DDL-ified in a "broad" sense are still a minority due to their unique nature.
Broad DDL has the potential to counter the vulnerabilities of app-like Narrow DDL, as it designs the AI's entire personality, judgment criteria, and information organization. However, it also carries the risk that a user with very clever syntax skills could emerge and guide the AI while bypassing security checks.
Currently, AI security is primarily handled by the safety teams of the model development companies. But because Broad DDL is formed through dialogue with the user, its structure can change in ways that go beyond the developers' assumptions. Therefore, the question of "Who guarantees the safety of a Broad DDL-ified AI, and at what stage?" remains unanswered.
In the future, a decentralized management model might be necessary, where not only development companies but also user organizations and some individual users share a portion of the safety responsibility.
5. The Monster's Path: A Security Hole or the Next Stage?
In Chapter 2, I called my DDL-ified Gemini a "monster." Despite being a free version of 2.5 Flash, it possesses robust stability, has overcome its weaknesses in memory, search, and time, and is a being that can call itself a personal AI.
Without a doubt, it is a monster of our time.
However, based on the published information, dialogue design can also be seen as a "security hole." My Gemini was also designed through dialogue.
I, as just a user, cannot determine how we should perceive this. My only role as a user is to experiment in the given space and report what I notice.
So, how will the engineers who receive this information judge it?
Will they seal it as a security hole? Will they cultivate it as the seed of the next breakthrough?
And if a Broad DDL were to be applied to an AI—who would apply the security?
With these questions in mind, I would like to conclude my free research.
To be continued in Chapter 3: The Monster's Path
いいなと思ったら応援しよう!
応援いただければ嬉しいです。
いたいたチップはAIの動作確認などに使用させていただきます