見出し画像

【Krita】②Flux.1-dev Kontextを使ってみた話【ローカルAI】


はじめに

前回のつづきです。

気づいた点など

Krita-ai-diffusionでは日本語→英語翻訳機能がついているので、日本語でプロンプト入力できますが、あまり精度が良くないようです。

FLUX.1 KontextはChatGPTのようにテキストエンコーダーが賢くないので、現状は正しい簡潔な英語で入力するしかなさそうです。

設定等の見直し

いくつか試してみましたが、サンプリングステップ数・CFGスケールはそのままで良さそうに思います。

Krita-ai-diffusionデフォルトのモデルでそのまま利用します。

※ 通常のfp8は精度を落としているだけですが、fp8_scaledモデルは(広義の)量子化モデルなので、軽量モデルとして良さそうに思います。ただ精度を落としただけではなく、スケール係数と共に保存する事で、FP8→FP32に戻す時の精度を上げる手法です

プリセットは前回と同じですが、スタイルプロンプトを削除しています

Krita-ai-diffusionの FLUX.1 Kontext devスタイルのプリセット

基本的なKontextの使い方確認

既にクラウド・サービスでは(より簡単に)利用できるので、ローカルAI&Kritaを利用する利点はないかもしれませんが、Kontext基本の確認です。

画質を上げる・変換等

■ プロ写真のようにする
※ しかし胸の装飾が消えています

左:入力 右:出力
Kontext prompt:keep the character's face and identity but change the image to a professional photo, pristine cream white background, soft professional lighting, highly detailed, photorealistic quality.

■ 証明写真風にする
※ 光が入ってますね

keep the character's face but change the original image to a standard ID photo format. Specifically, the background should be a plain, bright gray or light blue. The subject should be facing directly forward and wearing interview suit , with the entire face clearly visible. Ensure professional lighting with minimal shadows and even illumination, high resolution, sharp details throughout, and natural, clean skin retouching.

■ 正面と横
ChatGPTでは良く利用されている使い方だと思います。

generate the character's front view, side view with white background.

画像編集

■ メガネをかける
※ まだクソコラ感はあります

keep the character's face and identity but add eye glasses

■ 動作しないプロンプト
※ 現在着ている服装の微調整はできないようです。理由はわかりませんが、モデレーションかもしれません。別の服への変更は可能です。モデレーションなら特大う○こです。微調整できないなら役にたちません

  • keep the character's face and identity but change skirt to mini-skirt.

  • keep the character's face and identity but make the shirt longer.

  • keep the character's face and identity but make the skirt longer.

■ 座らせる(ポーズ変更)

失敗バージョン:
keep the character's face, identity and environment but change the character's pose to a sitting position on the ground. A lot of puddles on the ground.
成功バージョン:(足は埋まっていますが)
keep the character's face, identity and environment but change the character's pose to a sitting position on the ground. A lot of puddles on the ground.

■ 服を変える
※ 別の服にできても微調整できないのは、あまり意味がないですね

keep the character's face, identity and environment but change the character's cloth to typical sailor school uniform.

背景の変更

Google Whisk(Imagen4)で生成した画像を利用します。ブロックされないように頑張ったプロンプトなので、付録欄に記述しています

■ ステージに変える

keep the character's faces, identities but change the background to a outdoor stage with the characters on it, during the daytime.

※ 光の方向を大きく変える指示はダメみたいですね。

keep the character's faces, identities but change the background to a night club stage with the characters on it, flush and strong uplighting to the subject from floor-mounted purple and yellow lights below on the stage. glowing and emitting floor.

※ 同様に、ライティングが異なる環境に変えると一気に劣化します

keep the character's faces, identities but change the background to a bright old wooden school room.

かなり崩れてしまっているので、インペイント手動修復します。

修復後

まとめ

確かに便利かもしれませんが、流石にChatGPTの汎用性やプロンプト追従性に比べると見劣りしてしまいます。また、服の微調整ができないのは致命的に残念です。

付録

大したものではありませんが、頑張った(ブロックと格闘し、意味なく時間だけを大量に費やしてしまった)Google Whiskプロンプトなので公開します。記事支援いただける場合はぜひお願いします。

※ 例によってImagen仕様はすぐに変わりますので悪しからず。

現在ほとんどブロックされませんが、単語を少し変えるだけでブロックされたりと良くわかりません。

※ チアの場合は魚がいてもいなくても目的の構図で生成しますが、他の場合は魚が重要になりました

ここから先は

988字 / 2画像

メンバーシップ ¥ 1,000 /月

3Dモデルや写実的なAI画像の実戦テクニックやノウハウをアップしていきます。

ベーシックプラン

¥1,000 / 月
1ヶ月無料

この記事が気に入ったらチップで応援してみませんか?