見出し画像

ComfyUIでFP8/GGUF版のHunyuanImage-2.1を試す(推奨VRAM 12~16GB)

Last update 9-15-2025
※ 公開されて間もないため、何度か記事の修正等を行う場合があります。
高速な生成ができるDistilled版の記事もあります。
※ 関連記事は、ComfyUIの記事一覧にまとめてあります。




■ 0. 概要

▼ 0-0. はじめに

 本記事では、Windows上のComfyUIで「HunyuanImage-2.1」を利用するための手順を説明します。モデルはFP8版とGGUF版を利用します。

 それぞれの説明は下記の項目にあります。

  • 1. 留意事項と使用リソース量

  • 2.~3. FP8版の準備と生成

  • 4.~5. GGUF版の準備と生成

 Distilledモデルを利用した高速な生成手順は、下記の記事を参照してください。

 ComfyUIのインストール方法は、下記の記事をご覧ください。


▼ 0-1. HunyuanImage-2.1について

 HunyuanImageは、Tencentが開発しているHunyuanファミリーの一つに数えられる画像生成モデルで、2.1は9-9-2025に発表されたばかりです。

 特徴としては、オプションでPromptEnhancerとRefinerが存在すること、圧縮率が32倍のVAEを用いて4MP(2048x2048)の出力ができること、中国語と英語のテキストを描画できること、等があります。ある程度の日本語も描画できるようです。

 Text EncoderはQwen-Imageと同じ「Qwen2.5-VL-7B-Instruct」のほか、Googleの小型な「ByT5 - Small」を採用しています。モデル自体は英語や中国語のプロンプトに特化しているとみられますが、日本語もある程度は通ると考えられます。

 ライセンスは「TENCENT HUNYUAN COMMUNITY LICENSE」が採用されています。

 パラメーター数は17Bです。参考まで、SDXLが3.5B、SD 3.5 Largeが8B、FLUX.1が12B、HiDream-I1が17B、Qwen-Imageが20Bとなっています。

 下記はComfy Org(ComfyUIの開発元)の情報です。blogやドキュメントが追加されたら随時掲載します。


▼ 0-2. 関連リンク1

 無料で生成のおためしができます。

 生成AIプレイグラウンドのfalも利用可能です。1回あたり$0.1となっています。なお、価格は変動する場合があります。


▼ 0-3. 関連リンク2

 公式の各種モデル(Model、Text Encoder、VAE)です。本記事では利用しません。

 本記事で利用する、FP8版のModel、Text EncoderとVAEです。

 FP8版のModelはCivitaiにも掲載されています。Civitaiの仕組み上、ファイル名が異なる可能性があります。

 本記事で利用する、GGUF版の各種モデルと拡張機能です。



■ 1. 留意事項と使用リソース量

▼ 1-1. 留意事項

 ComfyUIを更新していないと、必要なノードが存在しない場合があります。バグフィックスを含むこともあるので、なるべく最新版にしてください。

 次回の生成を高速化するため、VRAMに読み込んだモデルが自動的にメインRAMへ待避(オフロード)されることがあります。そのため、メインRAMが32GB以上であることが望ましいです。


▼ 1-2. 生成時間

 筆者の環境(GeForce RTX 5060Ti 16GB)での生成時間は下記のとおりです。ディスクからのモデルの読み込みや、プロンプトの処理の時間は含まれていません。GGUF形式はModelのサイズ(量子化タイプ)によって多少変化します。

  • FP8
    115秒程度(20Steps、euler/simple、CFG=3.5)
    80秒程度(変更:weight_dtype=fp8_e4m3fn_fast)

  • GGUF Q4_K_M
    135秒程度(20Steps、euler/simple、CFG=3.5)

 Distilled版のモデルを利用すると、設定にもよりますが57~21秒程度まで下げることができました。



■ 2. 準備(FP8版)

▼ 2-1. はじめに

 必要な情報は下記のURLに記載されています。ただし、現時点では別のFP8版Modelを利用します。

掲載されているワークフロー

 下記は、ComfyUIがHunyuanImage-2.1をサポートした当初の情報です。

掲載されているワークフロー(WIP)

▼ 2-2. モデルの設置

 筆者は種類ごとにディレクトリを分けているのでそのように記述しますが、お好みで決めていただいて構いません。

 下記のURLに掲載されたそれぞれのファイルをダウンロードして、適切なフォルダへ移動します。設置済みの場合はスキップしてください。なお、Qwen-Imageの記事で利用したものと同一のText Encoderが含まれています。

  • Comfy-Org/HunyuanImage_2.1_ComfyUI
    https://huggingface.co/Comfy-Org/HunyuanImage_2.1_ComfyUI

    • split_files/text_encoders/
      qwen_2.5_vl_7b_fp8_scaled.safetensors (8.73GB)
      byt5_small_glyphxl_fp16.safetensors (418MB)
      → models\text_encoders\hunyuan_image\

    • split_files/vae/
      hunyuan_image_2.1_vae_fp16.safetensors (773MB)
      → models\vae\hunyuan_image\



■ 3. 生成(FP8版)

▼ 3-1. 概要

 生成用のワークフローを提供しますので、改造、再公開を含め自由にご利用ください。


▼ 3-2. 補足

 Modelのdtypeを「fp8_e4m3fn_fast」に変更すると、生成速度が上がります。品質が少し落ちるようなので、お好みで選んでください。


▼ 3-3. ワークフローと生成例

 下記のファイルをダウンロードして、ComfyUIの画面にドラッグ&ドロップしてください。

 ワークフローの全体です。プロンプトは、既存のものをClaudeとGrokで変更しました。筆者の環境では1回あたり135秒程度かかります。

ワークフローの全体

 解像度は、公式には「2048 x 2048 (1:1)」「2304 x 1792 (4:3)」「1792 x 2304 (3:4)」「2560 x 1536 (16:9)」「1536 x 2560 (9:16)」の5種類が挙げられています。Stepsは20、CFGは3.5です。

Japanese anime-style illustration of a young girl with iconic anime features: large, expressive brown eyes, a soft, rounded face, and long, arranged dark brown hair. She displays a shy, bashful smile with a subtle blush on her cheeks. Her outfit consists of a white short-sleeved blouse with a delicate lace collar, a navy floral midi skirt, white socks, and brown loafers. She sits at a rustic wooden table on a charming countryside cafe terrace, striking a modest pose with one hand resting against her cheek. A coffee cup sits on the table in front of her. A speech bubble is positioned above her head, containing the Japanese text "一緒に休も?" with a line break as specified. The background captures a serene, remote mountain village with quaint stone cottages, winding dirt paths, vibrant wildflower meadows, distant rolling hills, and grazing sheep. Warm golden hour sunlight filters through lush grape vines, casting dappled lighting across the scene. The peaceful rural atmosphere is enhanced by a soft earth-tone color palette, rendered in a detailed and vibrant anime art style.



■ 4. 準備(GGUF版)

▼ 4-1. はじめに

 FP8版と同じ流れで説明するので、内容が重複する部分があります。

 HunyuanImage Distilledと通常版の違いは、Modelと一部の設定のみです。また、Qwen-Imageと同じText Encoderを利用します。


▼ 4-2. モデルの設置

 筆者は種類ごとにディレクトリを分けているのでそのように記述しますが、お好みで決めていただいて構いません。

 下記のURLに掲載されたそれぞれのファイルをダウンロードして、適切なフォルダへ移動します。設置済みの場合はスキップしてください。ModelはQ4_K_Mでおよそ十分な品質となりますが、VRAM使用量と品質を見ながら上げ下げすることができます。なお、Qwen-Imageの記事で利用したものと同一のText Encoderが含まれています。


▼ 4-3. 必要な拡張機能について

 本記事に掲載したワークフローを読み込んだとき、拡張機能が不足していると「Some Nodes Are Missing」の表示が出ます。その場合は、ComfyUI Managerの機能を利用してインストールすることができます。

 拡張機能をインストールする方法は、下記の記事の「自動インストールの手順 」または「手動インストールの手順」を参照してください。

 本記事のワークフローでは下記の拡張機能を利用します。



■ 5. 生成(GGUF版)

▼ 5-1. 概要

 FP8版と同じ流れで説明するので、内容が重複する部分があります。

 生成用のワークフローを提供しますので、改造、再公開を含め自由にご利用ください。


▼ 5-2. 補足

 利用にあたり注意点があります。2回目以降でプロンプトを変更した場合は生成開始前に「Unload Models」を実行してください。専用GPUメモリ(VRAM)にデータが残った状態のままText Encoderのロードが行われて共有GPUメモリにはみ出し、プロンプトの処理時間がかなり長くなる場合があります。

適宜「Unload Models」を実行する

▼ 5-3. ワークフローと生成例

 下記のファイルをダウンロードして、ComfyUIの画面にドラッグ&ドロップしてください。

 ワークフローの全体です。プロンプトは、既存のものをClaudeとGrokで変更しました。筆者の環境では1回あたり135秒程度かかります。

ワークフローの全体

 解像度は、公式には「2048 x 2048 (1:1)」「2304 x 1792 (4:3)」「1792 x 2304 (3:4)」「2560 x 1536 (16:9)」「1536 x 2560 (9:16)」の5種類が挙げられています。Stepsは20、CFGは3.5です。

Japanese anime-style illustration of a young girl with iconic anime features: large, expressive brown eyes, a soft, rounded face, and long, arranged dark brown hair. She displays a shy, bashful smile with a subtle blush on her cheeks. Her outfit consists of a white short-sleeved blouse with a delicate lace collar, a navy floral midi skirt, white socks, and brown loafers. She sits at a rustic wooden table on a charming countryside cafe terrace, striking a modest pose with one hand resting against her cheek. A coffee cup sits on the table in front of her. A speech bubble is positioned above her head, containing the Japanese text "一緒に休も?" with a line break as specified. The background captures a serene, remote mountain village with quaint stone cottages, winding dirt paths, vibrant wildflower meadows, distant rolling hills, and grazing sheep. Warm golden hour sunlight filters through lush grape vines, casting dappled lighting across the scene. The peaceful rural atmosphere is enhanced by a soft earth-tone color palette, rendered in a detailed and vibrant anime art style.



■ 6. おまけ

▼ 6-1. おまけ画像

 HunyuanImage-2.1で生成した画像を何枚か掲載します。プロンプトはHunyuan-PromptEnhancerを通した後のものです。

An anime-style close-up portrait captures a joyful, shy Japanese girl looking directly at the viewer. Her face is framed by long, flowing silver-gray hair, which is meticulously styled with small double chignons, each tied with a vibrant red ribbon. She wears a flowing white dress and is elegantly playing a violin, with the instrument positioned gracefully near her shoulder. Shimmering musical notes float ethereally around her, conveying a sense of serene performance. The scene is set at night under a deep blue sky, dominated by a large, luminous full moon that illuminates the entire composition. Scattered stars dot the darkness above distant mountains, which are silhouetted in the background. The soft moonlight casts a gentle glow on her figure and the immediate surroundings, which include Japanese silver grass (susuki) with their distinctive bushy seed heads. A subtle wind creates gentle movement, causing her hair and the fabric of her dress to flow softly. The overall image is rendered in a detailed and emotive anime art style.
A vibrant, anime-style illustration captures a crisp autumn day in a park filled with maple trees. In the foreground stands a cheerful little girl with fair skin and big, expressive eyes. Her auburn hair is styled in intricate braids. She wears a cream-colored knit sweater sweater dress, complemented by a bright red scarf wrapped around her neck. With a joyful expression, she holds a comically large wooden sign, on which the text "Welcome to HunyuanImage-2.1!" is written in an adorable, wobbly script. A cascade of colorful maple leaves falls from its sides, appearing as if caught in a gentle breeze. Behind her, a winding stone path is covered in a thick layer of fallen red and gold leaves. In the middle distance, other visitors are depicted taking photographs, their figures slightly blurred to create a sense of depth. The background is dominated by a dense grove of maple trees, their brilliant red and gold foliage forming a rich canopy. The entire scene is illuminated by warm, golden afternoon sunlight filtering through the leaves, casting long shadows and creating a magical, glowing atmosphere with visible light particles dancing in the air. The artwork is presented as a high-quality digital illustration in a classic anime style.
A young girl with gray hair styled in pigtails sits attentively at a grand piano inside an ornate, dimly lit concert hall. The central figure is the girl, who wears a white ribbon blouse and a sky blue skirt detailed with small white floral patterns. Her delicate fingers are positioned on the ivory keys of the piano, suggesting she is about to play. Perched atop the polished black surface of the grand piano is an elegant orange tabby cat, its coat showing distinct stripes. In the background, the vast hall is defined by a high, vaulted ceiling, which is illuminated by large crystal chandeliers casting a warm, golden glow. Far below, the audience occupies the rows of deep red velvet seats, their forms rendered as dim silhouettes in the hall's recesses of light. The overall artistic style is a soft watercolor painting, influenced by anime character design, notable for its large, expressive eyes and delicate line work, which captures an intimate performance moment.
A young girl with smooth, fair skin is seated at a dessert buffet table within a bright and luxurious hotel restaurant. Her long, dark brown hair is styled in a high ponytail, accessorized with a simple red bow. Her facial features are characterized by large, warm brown eyes, a small nose, and a gentle, natural smile. She is wearing a pastel pink summer dress featuring delicate white lace trim along the neckline. In front of her, a white ceramic plate holds a modest assortment of desserts, including a slice of strawberry shortcake with visible layers, a small scoop of rich chocolate mousse, a couple of colorful fruit tarts, and a few pastel-colored macarons. Adjacent to her plate, a tall, clear glass contains iced tea, complete with visible ice cubes and a slice of lemon submerged within the liquid. The background is softly blurred, showing the indistinct forms of other dessert stations and the upscale interior of the restaurant, creating a shallow depth of field. The entire scene is bathed in soft, warm lighting that casts gentle highlights and creates a cozy, inviting atmosphere. This image presents a photorealistic photography style.
A slightly elevated, close-up shot captures a young Japanese woman in mid-stride while hiking on a sun-drenched mountain trail. The camera is positioned low and to the side, creating a dynamic diagonal perspective that focuses on her upper body and walking motion. She has joyful and energetic facial expressions, with her eyes looking forward, and her hair is pulled back into a high ponytail, partially covered by a simple cap or headband. She is dressed in light summer hiking gear, consisting of a colorful, short-sleeved tank top or a light, breathable hiking shirt, paired with shorts. In one hand, she grips a trekking pole, using it for balance as she moves along the path. A smaller daypack is visible, strapped to her side. The trail itself is narrow and unpaved, flanked by the lush, vibrant green vegetation of a mountainous region during summer. The background shows layers of forested slopes and distant mountain peaks under a bright, clear sky, with the entire scene illuminated by bright, natural sunlight. This image presents a photography style, characterized by a shallow depth of field that keeps the subject sharp against a softly blurred background.
A slightly elevated, close-up shot captures a professional businessman in mid-stride as he walks through a bustling New York City street, conveying a sense of dynamic energy and purpose. The central figure is a man with a determined expression, his face showing confidence as he moves forward. He is dressed in a sharp, well-tailored business suit of a deep navy or charcoal gray color, worn over a crisp dress shirt and a complementary tie. In one hand, he carries a dark leather briefcase or a sleek laptop bag. The camera angle is a low diagonal, focusing on his upper body and capturing his dynamic walking motion. He moves along a wide city sidewalk, where the pavement is lightly textured and reflective. In the background, the iconic urban scenery of New York City is rendered with a shallow depth of field, creating a soft, blurred bokeh effect. Recognizable shapes include towering skyscrapers, the vibrant yellow of a taxi cabs, and the general bustle of city life. Natural daytime lighting illuminates the scene, casting soft shadows and highlighting the textures of his suit and the surrounding environment. This image presents a photography style.
A lively scene depicts a group of diverse plush dolls engaged in conversation within a brightly lit school hallway. In the foreground, three central dolls form an animated group. The doll on the left is a girl with long brown hair styled in pigtails, wearing a classic navy blue and white sailor-style uniform with a red ribbon. She holds a miniature textbook in one hand and a tiny pencil case in the other, her expression animated as if speaking. The central doll has short black hair and wears a light pink sailor uniform with a matching bow, holding a cute pink eraser and tilting her head curiously. To the right, a third doll with blonde hair in a neat bun is dressed in a blue-and-white uniform with a yellow ribbon, gesturing with one hand as she listens. The background features the hallway setting, including a row of colorful lockers lining one wall and large windows on the other, which allow warm, cheerful light to stream in. A bulletin board, scaled down to fit the doll's height, is mounted on the wall between the lockers. The entire scene is rendered in a soft-focus, 3D digital art style, emphasizing the plush texture and fuzzy details of the toys.
A surreal and cinematic landscape unfolds on an alien world, where realism and fantasy converge. In the foreground, unknown beings with elongated limbs and skin that has a glossy, liquid-like quality move gracefully through the dreamscape. They glide over shimmering metallic vegetation that shifts between organic, leaf-like plants and sharp, geometric crystal formations. The middle ground is dominated by a series of massive, spiraling towers that ascend towards the sky, their metallic surfaces twisting into complex, fantastical shapes. Above the towers, strange bioluminescent creatures, resembling living jellyfish with translucent, bell-shaped bodies and long, delicate tentacles, drift slowly through the air, emitting a soft, ethereal glow. The sky itself is a deep, cosmic expanse filled with multiple moons of varying sizes and colors, which are surrounded by vibrant aurora streams that swirl in shades of green, purple, and blue. The overall lighting shifts between warm, golden ethereal glows and cool, blue-toned illuminations, casting dramatic shadows and highlights across the entire scene. This image is rendered in a highly detailed, surreal cinematic style, reminiscent of a still from an epic science-fiction film.



■ 7. その他

 私が書いた他の記事は、メニューよりたどってください。

 ComfyUIに限定した記事の一覧もあります。

 記事に関することで何かありましたら、Xの@riddi0908までお願いします。

いいなと思ったら応援しよう!