見出し画像

【ズボラ】AIが優秀すぎて、どんどんプロンプトを気にしなくてもよくなってきてると思う話


はじめに

規模が小さなAIであれば、その内部技術的な仕組みに沿ったプロンプトや効率的な使い方で大きな効果を期待できますが、規模が大きく賢くなればなるほど、利用において技術的な手法の意味が薄れます。

そして、小規模なAIや、特定用途に特化した制約・技術を要するAIほど、より高度なAIに条件を理解させ、適切なプロンプトを作成できる事を鑑みると、人間の作る文字表現力やプロンプト自体の(スキル)価値は、かなり低下しているように感じます。

個々のスキルの必要のない、全人類のスキルが均一化するAI社会に向かって進んでいるのかもしれません。それは「AI使いこなしスキル」とて同じです。AIを使いこなすためにAIを使うのが最も有効ですから。

筆者は、従来型の発明物であるパソコンやスマホ等の既存IT技術と異なり、AIの(使いこなし)スキルは、将来的に必要無くなるのではと考えています。

むしろ、探究心・強い意思・欲望のような生物の主体として原始的な原動力や行動力こそが、人類社会の競争に最も必要なものになるのではと思います。

本記事では、ズボラを極めるため、各画像生成AIのルールをAIに調査させ、それをコピペ・メタプロンプトとした例を紹介します。

そもそもAIは線形ではない

重みなどのAIネットワークは線形(線形行列で処理)ですが、プロンプトなどの入出力は線形ではありません。つまり、プロンプトをトークンに分けて古典的に(階層・線形)分類する手法は、大規模なものでは既に意味を失っています。

性能の高いAIであればあるほど「褒めたり丁寧な言葉でAIに接すればAIの性能が上がる」というような、一見すると宗教的(人間的)に思えてしまうテクニックの方がより適切なものになりつつあります。

AIが高度化するにつれて、その傾向はさらに強まるものと考えられます。

そして、逆を言うと、技術的に正しく詳細で複雑な説明が、投資詐欺のような解説になってしまいかねない状況です。各論として技術的に間違っていなくても、マクロ・統計・カオスのさざ波に消えてしまうような効果の可能性があります。

良く技術的根拠から、AI用プロンプトは〜に注意して作る「べき」と言われますが、そのような明確なものがあるのなら、それこそAIにさせるべきです。人間はさらにズボラを極めた褒めるだけのようなメタの世界に行く「べき」です。

※ ただし、Stable Diffusionなどの昔の規模の小さなものは大きな意味があります(プロンプト集、ネガティブプロンプトや層の指定が有効な手法になります)

※ LLMにおいても、現在のAI性能では、明確で簡潔な指示の方が良い結果を得られる事も事実です。特に、何度もやり取りするのではなく、最初の指示に一球入魂した方が良い性能が出るのも事実です。しかし、将来的にそれは無くなる傾向であり、むしろ人間と同じで、AGIに近づくにつれて、余計かつ無駄なノイズ情報があった方がより良い回答になるのではと筆者は考えています

FLUX.1 Kontextの例

※ FLUX.1-dev Kontextは、ローカルで自然言語指示による画像編集を可能にするモデルです。

筆者も最初のうちは、公式サイトのプロンプト例や作り方解説を参考にして自分で考えていましたが、AIに任せた方が成功する確率が高い事に気が付きました。

Kontext用のプロンプトを生成するメタプロンプトは、Geminiで作ったものですが、プロンプトを作る上でのルールを列挙させたにすぎません。大きく変更することなく、そのままルールをコピペするだけです。

予想以上に良い結果を出すので、別の画像生成モデルでも試しました。

Stable Diffusion 1.5の例

自然言語での記述に適さない小規模なAI例ですが、プロンプトの作り方が非常に機械的で、明確なルールが確立しているモデルです。

まずは、それらルールをGemini等のネット情報にアクセスできるAIに列挙させます。

※ 最新情報が必要なものはローカルAIでは力不足です。もちろんネット検索させれば可能ですが、それは完全なローカルAIではありません。

Gemini 2.5 Flash

この結果を「以下、注意点」としてコピペするだけでも、それなりのメタプロンプトになります。

重要な事は、それらにテクニックや作業がいらない事です。確認なしの「コピペ」だけです。

※ ただし、付録欄に添付したものは筆者が少し手を加えたものです。明らかに無駄な情報を削除・整形しています。

「花壇で屈んで座っている少女の画像」を生成してとします。

ローカルのQwen3 8B

(masterpiece, best quality, ultra detailed:1.2), a young girl, (18 years old:0.8), long flowing auburn hair, (bright green eyes:1.1), fair skin, wearing a (white cotton dress:1.0), kneeling in a flowerbed, (sunlit:1.3), (soft lighting:0.9), vibrant colors, (detailed flowers:1.2), (bokeh:0.7), cinematic composition, (shallow depth of field:1.1), (intricate details), (photorealistic:1.1)

Negative prompt: (ugly, distorted, blurry:1.3), low quality, bad anatomy, deformed, disfigured, watermark, text, signature, extra fingers, mutated hands, poorly drawn hands, unrealistic, cartoon, anime

生成されたプロンプト
SD1.5画像 ①

「80年代映画に出てくるような荘厳な機関車」でプロンプトを作ります

SD1.5画像②

SDXLの例

まったく同じ手法を利用します。

Geminiで作る

上記結果を注意点として、「花壇で屈んで座っている少女の画像」を生成してとします。

A young girl with long flowing dark hair, wearing a white summer dress, sitting on the ground in a flower bed, bent slightly forward with her arms resting on her knees, looking at the flowers, soft sunlight filtering through the trees, a peaceful and serene atmosphere, natural lighting, high detail, realistic style, photorealistic, cinematic composition, wide-angle shot, vibrant flowers in full bloom, soft shadows, 8K resolution, ultra-detailed, high-quality, artistic photography, natural textures, soft focus on the background, sharp focus on the girl's face and hands

Positiveプロンプト

ugly, deformed, blurry, low quality, bad anatomy, extra limbs, fewer limbs, out of frame, disfigured, poor lighting, grayscale, low resolution, unrealistic proportions, cartoonish, anime style, unrealistic colors, over-saturated, over-exposed, under-exposed, grainy, noise, artifacts, low detail, unrealistic textures, poor composition

Negativeプロンプト
SDXLで生成した画像①

同様に、「80年代映画に出てくるような荘厳な機関車」で生成したプロンプトです。

A majestic, ornate steam locomotive from the 1980s movie era, with brass fittings, polished metal, and intricate engravings, standing on a vast, dusty railway platform at dusk, surrounded by old-fashioned signal lights, steam rising from the engine, glowing red and orange lights, dramatic lighting, cinematic atmosphere, wide-angle shot, high detail, photorealistic, 80s movie style, retro industrial design, powerful and imposing, with a sense of grandeur and nostalgia, golden hour lighting, rich textures, deep shadows, cinematic composition, rule of thirds, high resolution, 8k detail

Positiveプロンプト

low quality, blurry, deformed, extra limbs, missing limbs, out of frame, disfigured, poor lighting, grayscale, unrealistic proportions, cartoonish, flat colors, low resolution, lack of detail, modern design, futuristic elements, digital artifacts, unrealistic steam, unrealistic lighting, over-saturated colors, unrealistic textures

Negativeプロンプト
SDXLで生成した画像②

FLUX.1の例

理由はわかりませんが、FLUX.1の情報の多くは(Geminiで)ブロックされました。仕方がないのでディープリサーチで試しました。

※ やはり人間思考が求められるのは、モデレーション(?)が発生した時でしょう

Geminiディープリサーチ
(指示と異なり)広範囲な内容のレポートになっているので、プロンプト部のみをコピペ利用します

「花壇で屈んで座っている少女の画像」として生成します。

A young girl, approximately 12 years old, is sitting cross-legged in a flower bed, bent slightly forward as if observing the ground closely. Her long, wavy chestnut hair is slightly tousled, with strands framing her face. She is wearing a light blue sweater and a white t-shirt, with a small yellow sunhat on her head. The sunlight filters through the leaves above, casting soft, warm light on her face and the flowers around her. The setting is a well-maintained garden with vibrant flowers like daisies, tulips, and lavender, with a gentle breeze causing slight movement in the petals. The scene is captured with a Canon EOS R5 with a 85mm f/1.8 lens at f/2.8, creating a soft bokeh effect that emphasizes the girl and the flowers. The lighting is golden hour, with a warm glow that adds depth and mood to the image. The atmosphere is peaceful and serene, evoking a sense of quiet observation and connection with nature.

FLUX.1 devで生成した画像①

「80年代映画に出てくるような荘厳な機関車」で生成したプロンプトです。

A grand, steam-powered locomotive in the style of 1980s cinema, captured with a high-end Hasselblad X2D 100C camera using an XCD 90V lens at f/4. The locomotive is massive, with a dark, polished brass and iron exterior, intricate rivet patterns, and a towering smokestack that emits thick, swirling steam. The engine compartment is open, revealing complex pistons, glowing red-hot cylinders, and a massive flywheel. The locomotive is set on a gravelly, industrial track at dawn, with soft morning light casting long shadows and highlighting the texture of the metal. The atmosphere is dramatic, with a sense of power and nostalgia, evoking the grandeur of a classic train from a 1980s epic film. The image should have a cinematic feel with rich color grading, deep contrast, and a sense of scale and movement.

FLUX.1 devで生成した画像②

Google Imagen3/4の例

ローカルAI生成ではありませんが、日本的な画像生成では圧倒的な性能を誇るImagen4で試します。

ImagenもFlux.1と同様に情報を絞られました。ディープリサーチ結果です。

プロンプト部のみを抜き出して、メタプロンプトとして利用します。

「花壇で屈んで座っている少女の画像」

予想通りの展開です←これを言いたくて"屈み"ました

人間脳の登場です。22歳ぐらいにしておけば大丈夫でしょう。

つまり、モデレーションがある画像生成AIは、正攻法・正しい手法しかできないAIではダメという事です。独善的な事しか出来ない検閲AIと相性が悪いです。

※ メタな作業として人類に残された仕事は、AIに禁じられた処理を扱う事のみです。

Whisk(Imagen4)で生成した画像①

「80年代映画に出てくるような荘厳な機関車」

Whisk(Imagen4)で生成した画像②

まとめ

FLUX.1やImagenの「ルール」取得が困難だった理由の一つは、比較的新しいという事もありますが、テキストを自然言語としてまともに扱えるため、Stable Diffusionのように機械的で確定したやり方や信頼できる手法が見つかりにくかったのでしょう。ネット上の各所で全く逆のやり方を推奨していれば、AIは信頼できない情報として扱ってしまったのかもしれません。

本記事の手法は、画像生成に限らず、ほぼすべてにおいて有効だと思います。何かを作る際には、必ず「あーしろ・こーしろ」推奨ルールがあります。

もし、その推奨ルールをAIに食べてもらったにも関わらず、想定したものができない場合は、(AIの推論力が優秀であると仮定すれば)推奨ルールが間違っている・矛盾をはらんでいるという話になります。

そして、技術に基づいた・確定した推奨ルールがない場合は、AIと自然言語のやり取りを通じて自由に「創造」する事こそ、現代のAI活用の醍醐味だと思います。

付録

本記事で作成したメタプロンプトを、Open WebUIのプロンプトショートカットにしたものを添付します。例によってほとんどAI生成そのままで、精査できておらず、支援者様(メンバーシップ)専用にさせていただきます。参考程度にご利用ください。

極限まで練り上げられたプロンプトでなくとも、大雑把なコピペ程度で良い性能が出せるという例です。ただし、コンテキストサイズは大きくなりますが。

ここから先は

24,992字

メンバーシップ ¥ 1,000 /月

3Dモデルや写実的なAI画像の実戦テクニックやノウハウをアップしていきます。

ベーシックプラン

¥1,000 / 月
1ヶ月無料

この記事が気に入ったらチップで応援してみませんか?