anima_pencil-XL_clear を公開! 自然言語のプロンプトとクリアーな出力 【画像生成 AI・SDXL】
はじめに
こんにちは、きまま / Easygoing です。
今回は、自然言語入力が使えてクリアーなイラストを出力できる anima_pencil-XL_clear モデルを公開したのでご紹介します。

anima_pencil-XL モデルって何?
anima_pencil-XL モデルは、ぶるぺん さんが公開されている SDXL のアニメモデルです。

ぶるぺん さんの blue_pencil-XL モデルは、実写からアニメ系まで膨大な数のモデルをマージしたアニメモデルで、プロンプトを解析する CLIP を可能な限り巻き戻しているため、自然言語のプロンプト を理解できるのが最大の特徴です。
anima_pencil-XL ≒ blue_pencil-XL + Animagine-XL
そして、anima_pencil-XL モデルは、blue_pencil-XL モデル と Animagine-XL およびその派生モデルをマージして、 Animagine-XL シリーズの バリエーション も取り入れたモデルになっています。
anima_pencil-XL モデルの特徴!
それでは、anima_pencil-XL モデルと Animagine-XL モデルのイラストを並べてみます。

anima_pencil-XL シリーズの特徴
愛嬌のある丸顔・丸目
自然言語のプロンプトが使える
イラストが明るい
Animagine-XL や illustrious-XL をはじめ、多くのアニメモデルは 小顔・細目の整った顔 を描写しますが、anima_pencil-XL モデルは独特な 愛嬌のある丸顔 を出力します。
anima_pencil-XL モデルは、CLIP の安定性が高いので、masterpiece などの 品質プロンプト と ネガティブプロンプト は必須ではありません。
さらに、anima_pencil-XL モデルは 顔が常に明るく描写される ため、明るい雰囲気のイラストを作りやすくなっています。
anima_pencil-XL_clear の改良点
それでは、今回調整した anima_pencil-XL_clear モデルの改良点をご紹介します。
左側が今回調整した anima_pencil-XL_clear モデル、右側がオリジナルの anima_pencil-XL-v5.0.0 モデルの出力になります。
紫の魔女

ブレザーの女の子

拡大図

anima_pencil-XL-v5.0.0_clear モデルは、コントラストが高く、鮮やかな発色 になるように調整しています。
また、目元や衣類の拡大してみても、左の anima_pencil-XL-v.5.0.0_clear モデルの方が、クリアー な表現で ディティールも正確 になっています。
そのほかの作例!
続けて、そのほかの作例も見てみましょう。
カボチャの国

an animated female character, resembling a witch, donning a purple witch's hat with golden embellishments, stands against a darkened backdrop filled with levitating items like a candy cane and a carved pumpkin. her eyes are closed, and she wears a matching purple and black attire with a matching hat featuring a smiling face. the overall mood is mysterious and enchanting.,
dutch angle, dynamic angle, close up, upper body, happy, smile, laugh, peaceful, wind, vivid, colorful, small breasts, mature, elegant, black background, burgundy, rouge, alizarin, royal purple, indigo, deep blue
縁日

a young female character in a kimono with a red flower in her hair stands in a traditional japanese street at night, surrounded by lanterns, and people, with a warm and serene ambiance.,
dutch angle, dynamic angle, close up, upper body, happy, smile, laugh, peaceful, wind, vivid, colorful, small breasts, mature, elegant, black background, burgundy, rouge, alizarin, royal purple, indigo, deep blue
紅白のバラ

an animated female character with silver hair and blue eyes holds a bouquet of white roses in a room with a city view, wearing a red and black outfit with a black choker, surrounded by a warm and cozy atmosphere with soft lighting and a framed picture on the wall.,
dutch angle, dynamic angle, close up, upper body, small breasts, mature, elegant, black background, burgundy, rouge, alizarin, royal purple, indigo, deep blue
豹

a close - up of a leopard's face, with striking yellow eyes and detailed fur, is depicted against a dark background with a hint of blue and black colors. the leopard's fur is detailed, and its eyes are focused intently on the panther's face. the overall style is realistic and detailed, with a focus on the leopard's fur and the eyes.,
dutch angle, dynamic angle, close up, upper body, happy, smile, laugh, peaceful, wind, vivid, colorful, black background, burgundy, rouge, alizarin, royal purple, indigo, deep blue
鳥人

a vividly colored rooster stands prominently against a dark background, its vibrant plumage contrasting with the darker tones of the surrounding environment. its striking plumage stands out prominently, with its striking plumage accentuating its striking colors. the bird's talons are a mix of red, blue, and purple, adding a touch of surrealism to the scene.,
dutch angle, dynamic angle, close up, upper body, happy, smile, laugh, peaceful, wind, vivid, colorful, small breasts, mature, elegant, black background, burgundy, rouge, alizarin, royal purple, indigo, deep blue
anima_pencil-XL_clear のレシピ
anima_pencil-XL シリーズには、Fair AI Public License 1.0-SD ライセンスが設定されているため、派生モデルを公開するときは マージレシピの共有 が必要です。
anima_pencil-XL-v5.0.0_clear のマージレシピ

anima_pencil-XL_clear モデルは、オリジナルの anima_pencil-XL-v5.0.0 に対して UNET と VAE の調整を行っています。
UNET:過剰に反応している層(time_embed. / label_embed.)の抑制
VAE:アニメ専用 VAE(SDXL Anime VAE Dec-only B3)に換装
anima_pencil-XL_clear モデルでは、UNET の初期の層の中で 過剰に反応している層 の働きを抑えています。
また、VAE を アニメイラスト専用 のものに換装して、色の表現が鮮やかになるようにしています。
anima_pencil-XL_clear は容量が大きい?
anima_pencil-XL_clear モデルは、一般的な SDXL モデルと比べて少し容量が大きくなっています。
anima_pencil-XL_clear:6.61 GB
一般的な SDXL モデル:6.46 GB
その理由は、換装した VAE を FP32 形式 で結合した ためです。
VAE は BF16 形式で処理される
SDXL は FP16 形式で普及しましたが、初期に VAE の処理が不安定で黒い画像が生成される ことがありました。

そのため、ComfyUI / Stable Diffusion webUI ともに、現在は対応している GPU では VAE の処理を BF16 形式で行う ようになっています。
FP32 → BF16:劣化なし
FP16 → BF16:わずかに劣化する
このとき、FP32 → BF16 形式の変換であれば劣化は起こりませんが、FP16 → BF16 形式の変換だとわずかに劣化が生じてしまいます。

実際に目に見えるほどの画質の違いはありませんが、VRAM 使用量 と イラストの生成時間 は変わらないので、anima_pencil-XL-v5.0.0_clear では FP32 形式の VAE を採用しています。
CLIP-Lで自動プロンプト!
今回の自然言語のプロンプトは、comfy-cliption カスタムノードで CLIP-L を使って画像をキャプションして作りました。
入力画像

CLIP-L のキャプション
紫色の髪と青い目の若い女性キャラクターが、ふわふわのオレンジ色の猫(青い目)を抱いており、ピンクのセーターと青いオーバーオールを着て、紫と青の筋が入った暗い背景を背に、喜びと好奇心の感覚を伝えている。
CLIP-L は、前回使用した Florence-2 よりキャプションは不正確ですが、知識のベースが広いため 解釈の幅が広い のが特徴です。
CLIP-L は 色相 や 感情 まで言語化してくれるため、創造的なアート のイラストの生成に適しています。
また、自然言語のプロンプトはタグ付けのプロンプトより情報量が多いため、バリエーション や 繊細な表現 を求めるときにも最適です。
まとめ:anima_pencil-XL_clear を使ってみよう
anima_pencil-XL_clear を公開
自然言語のプロンプト
愛嬌のある丸顔 と 柔らかい表現
今回は、anima_pencil-XL-v5.0.0_clear モデルをご紹介しました。
anima_pencil-XL シリーズは、登場から既に1年半が経過していて、ディティールの正確さは最新のモデルに劣りますが、自然言語のプロンプト が使えて、柔らかな表現 ができるアニメモデルとしては、現在でも唯一無二のモデルになっています。

anima_pencil-XL モデルは使いやすく、今回の改良で 最新のモデルにも匹敵する表現力 を手に入れた思うので、皆さんも是非試してみて下さい。
最後までお読みいただきありがとうございます!
参考記事
文章 vs 単語 のプロンプト
kanesi さんが、自然言語とタグ入力のプロンプトについて比較されています。
実例付きでとても分かりやすいので、皆様も是非一度ご覧ください。
