Scribble at 2026-08-08 20:16:52 Last modified: 2026-08-09 12:24:58
A 35mm lens at f/8 captures the Victorian façade of London’s fog-draped streets, its ornate stonework rendered in precise grayscale tones from Zone III to Zone VI, where weathered brickwork absorbs ambient light while polished iron balconies reflect the overcast sky’s diffused glow. Cobblestone avenues, their surfaces etched with decades of wear, contrast with the smooth, reflective glass of shopfronts, their edges defined by sharp highlights and gradual shadow roll-off. A red double-decker bus, its paint dulled by soot, occupies the foreground, its silhouette framed by the stark geometry of narrow alleyways, their depths receding into Zone IV shadows. The iconic clock tower looms in the distance, its Gothic spires slicing through the low-hanging fog, while gas lamps cast elongated, creamy bokeh across the pavement, their light sources rendered as soft, luminous spheres. Reflective windows of adjacent buildings mirror the scene, doubling the architectural details in a symmetrical interplay of light and form. The overcast sky, a uniform Zone V gray, bathes the scene in even luminance, revealing micro-contrast in the texture of flower boxes and the intricate carvings of historic facades. A long exposure blurs the movement of unseen pedestrians, their silhouettes dissolving into the fog’s density, while the dynamic composition balances varied building heights, from the low-rise shops to the tower’s towering silhouette. The camera’s shallow depth of field isolates the focal plane—the wrought iron details of a doorway—against the broader urban sprawl, emphasizing scale through meticulous tonal gradation. The final image, a testament to the Zone System’s precision, echoes the documentary rigor of Ansel Adams, merging technical mastery with the unyielding clarity of silver gelatin’s grayscale reality.
Steps: 8, Sampler: DPM++ 2s a RF, Schedule type: Linear Quadratic, CFG scale: 1, Shift: 3, Seed: 4277665692, Size: 1376x768, Model: ernieTurboFP8_v10, Model hash: efb8e77e30, Module 1: lensVAEFLUX2_lensBundleBf16, Module 2: ministral-3-3b, RNG: CPU, Version: neo-2.26
Baidu から今年の4月にオープン・ウェイトとしてリリースされた ERNIE-Image Turbo という新しいモデルを使い始めた。Text to Image Leaderboard (Open Weights) at Artificial Analysis で、Ideogram などには及ばないものの、オープン・ウェイトだけの順位で言えば Qwen Image Max 2512 や FLUX.2 Klein 9B などよりも高い評価を得ているらしい。実際に上のような画像を出してみたが、詳細で長いプロンプトをしっかり書けば、それに見合う画像を生成してくれるようだ。ただし、長くなればなるほど整合的な内容でなくては簡単にガラクタとなってしまうため、最初から丁寧にプロンプトを作るか、Z-Image Engineer のようなプロンプトの整形用モデルを使う必要があるかもしれない。
ちなみに、ERNIE-Image シリーズはテキスト・エンハンサーとして Ministral 3B を使っているのだが、これをプロンプトの調整用にしたという "Prompt Enhancer" なるファイルとして置き換えて使うかのような扱いになっているのだが、これは全くの錯覚だ。Baidu がリリースしている公式の Prompt Enhancer(中身は、MInistral 3B のカスタム版である)を代わりに使ってみたが、プロンプトの書き換えなんて処理の途中で内部的にやってくれたりはしない。プロンプトの書き換えを専用のノードとして定義できる ComfyUI 専用のファイルかもしれないので、ひとまず Forge Neo のユーザは Prompt Enhancer の話は無視して Ministral 3B を使っておけばいいと思う。というか、どう考えても最初からまともなレベルのプロンプトを Prompt Enhancer などに頼らず自分で作ることが原則だろう。そのためだけでもいいから、せめて高校レベルまでの英語は勉強しておくべきだ。
なお、ファイル構成としては、本体の拡散モデルが約 17 GB、テキスト・エンハンサーが約 8 GB、そして VAE が約 350 MB となっているため、VRAM が 16 GB である僕のマシンでは必ずメイン・メモリへのオフロードが発生する。したがって、1,376 x 768 ピクセルの画像を1枚だけ生成するのに30秒もかかる。