Loading…
Loading…

創作幕後、工具分享與靈感隨筆
AI ImageQwen-Image-2.1 一個模型同時做「參考頭像變九種表情」與「原生透明背景」,成品直接是透明背景的 PNG,不需要再用去背工具處理。關鍵在 prompt 怎麼寫,不在 ComfyUI 的設定:只寫「透明背景」會得到九張不透明卡片,要明說每一格內人物與文字以外全部透明。這篇公開完整 prompt、六個角色的範例、十六宮格寫法,以及從生成到自動切格縮到 LINE 規格的 ComfyUI 工作流。
AI ImageQwen-Image-2.1 是阿里九月開源的生圖模型,一顆權重同時做文生圖與圖片編輯,主打原生透明背景輸出、最多十張參考圖、原生 2K。我用四章 182 張圖在自己的電腦上逐條對照:透明輸出在玻璃、煙、頭紗這類半透明題目確實是去背模型做不到的,但拿它當一般去背工具會輸給專用模型;單張參考圖的編輯能把沒被要求改的地方保持不動,多張參考圖的合成則會畫出分身;繁體中文 53 題 30 題勝出,與 SenseNova 並列目前最穩;原生 2K 十題全部贏過 1K 再放大。
AI MusicACE-Step 1.5、MiniMax Music 3、HeartMuLa 3B 與 Stable Audio 3 那篇實測用的全部題目:4 首古風、5 首西方民俗舞曲、2 首對照、6 段 BGM 與 12 個音效。每題附曲風規格、三種模型各自的 prompt 寫法與完整歌詞,拿去測你手上的模型。
AI MusicSuno 九月起 Pro 每月只能下載 20 首,我把 ACE-Step 1.5、MiniMax Music 3、HeartMuLa 3B 三個開源歌曲模型,加上做 BGM 與音效的 Stable Audio 3,全部裝進本機 ComfyUI,用同一份歌詞、同一組 seed 跑 17 首歌與 12 個音效。古風、匈牙利舞曲、俄羅斯舞曲、剪片用的鋪底與音效,全部放上來給你自己聽。
ComfyUI30 條去背路線同圖實測之後,把結果收成一份能直接照做的選擇指南:先問三個問題,再對照你的圖是髮絲、玻璃、文字 Logo、插畫還是商品圖,各有該用的模型。附反差色流程的正確開法、授權對照,以及 2026 年不用再裝的舊模型清單。
Tutorial整合六篇實測與 53 組題庫,回答一個問題:要在圖裡寫繁體中文,2026 年 9 月該選哪個開源模型。海報、菜單、密集小字、漫畫對白、生成後改字,各自有答案;也附上不換模型就能少錯字的三個做法,和一套可以自己跑的題庫與評分表。
AI ImageLens 是微軟六月開源的 3.8B 生圖模型,主打「描述寫得好比模型大更重要」,用一個 20B 的語言模型當文字編碼器,論文說遵循度贏過 Flux.2 Klein 與 Z-Image、文字渲染也贏。我在自己的電腦上把它點名的對手全部叫出來同題對照:遵循度贏過 Klein 是真的,贏過 Z-Image 則看不出差別;英文長文字反而輸給 Z-Image 與 Qwen;繁體中文不行。另外解開一個疑問:為什麼 4 步的 Turbo 版常常比 20 步的標準版好看。
AI ImageHiDream-O1 是五月開源的 8B 生圖模型,官方介紹寫了六件事:沒有 VAE 所以細節更多、長文字與多語言排版、指令編輯、多張參考圖合成、一次生成分鏡與 15 種鏡位、還有一個會先推理再改寫 prompt 的 Agent。我用一套全新的 63 題在本機逐條驗證:多參考圖合成與 Prompt Agent 是真的強,細節多但雜訊也多,繁體中文錯字率仍高,15 種鏡位大半看不出差別。
AI Image為了逐條驗證 HiDream-O1 官方宣稱而寫的 63 組全新 prompt 全部公開:質感特寫、繁體中文與英文長文、指令編輯與多參考圖合成、四格分鏡與 15 種鏡位、需要背景知識的短句,以及官方 Prompt Agent 把短句改寫後的完整版本。附中文說明,拿去測你手上的任何模型。
AI Image上一篇測 SenseNova U1.5 時留了一個伏筆:官方的 Prompt 增強配方值不值得多一道工序?這篇用同 seed A/B 對照驗證:16 組題目、兩輪測試,從資訊圖表、藥品仿單一路測到漫畫和人像。結論比預期更清楚——增強對資訊密集版面是大勝, 但對漫畫美感沒有加分,人像甚至反而是原始 prompt 比較好。順便看清楚了一件事:這次測試裡的亂碼都長在哪裡。
AI Image從海報、菜單、藥品說明書到毛筆書法,17 類 53 組專測繁體中文文字渲染的 prompt 全部公開,附中文說明。這套題目養出了「濱」「廳」「衛」等全模型死亡字,先後測倒了五個開源模型,直到 SenseNova U1.5 才破關。拿去測你手上的任何模型。
AI Image三個月前用 53 組 prompt 測了五個開源 T2I 模型的繁體中文,結論是「沒有一個能穩定過關」。商湯八月底開源的 SenseNova U1.5 用同一套題目重跑一遍:「濱」「廳」「衛」這些全模型死亡字全部寫對,還能只改海報上的文字、其他一個像素都不動。小字仍然會錯,但繁體中文生圖的天花板確實被抬高了。
AI Image17 組動漫 prompt、同一組 seed,實測 Anima 全家四個版本(含社群加深版 2.9B)、SDXL 動漫王者 NoobAI-XL、新架構 NetaYume Lumina、Z-Image-Anime,再拉泛用模型 Krea2 當對照組。打鬥誇飾、手部完整度、風格多樣性一次測完——動漫題到底該不該用動漫專精模型?
ComfyUI把 ComfyUI 裡找得到的放大模型與臉部修復模型全部接進同一張工作流,用 12 張有「標準答案」的測試圖(先生成高清原圖、再程式降質成 128/256/512 像素與仿老照片)加 3 張真老照片,同圖實測 37 條路線,除了看圖還算了「像不像原人」的分數。哪款該當日常主力、修臉為什麼會換臉、老照片該找誰、硬碟裡那一堆放大模型哪幾款值得留,這篇從放大原理講起,一次說清楚。
ComfyUI把 ComfyUI 裡找得到的去背模型全部接進同一張工作流,用 15 張專門找碴的圖(髮絲、玻璃、煙、頭紗、Logo、動漫、單色背景)同圖實測 30 條路線。哪一款該當日常主力、半透明要找誰、先把背景改成反差色再去背到底好不好用,這篇從去背原理講起,一次說清楚。
AI Image用 30 組動漫題、同一組 seed,實測 Mage-Flow、Krea2、Boogu、Z-Image、Qwen 2512、Ideogram 4 六家開源模型:王道人像、90 年代復古、恐怖漫畫、可愛吉祥物、2.5D 賽璐璐、水墨動畫全部下場,附一組「風格錨點」對照實驗。
AI Image五個開源圖像編輯模型、18 道編輯考題、同一組來源圖與 seed:移除物件、改視角、換季、改表情、多行英文、中文字全部實測。誰聽話、誰保真、誰的文字最漂亮,一次看完。
AI Image用 16 道編輯考題(替換、移除、換季、改視角、改文字⋯⋯)實測微軟 Mage-Flow Edit 的四個變體,再拉 10B 的 Boogu Edit 同場對照。4B 的小模型在圖像編輯上,指令執行力意外地兇。
AI Image用同一套 35 組 prompt、同一組 seed,實測微軟 Mage-Flow 的 Standard/Turbo × bf16/int8 四個變體,再拉 Krea2 Turbo 同場對照。4B 參數、MIT 授權的輕量模型,到底能打到什麼程度?
AI Image社群上有一堆號稱「讓 AI 圖更細緻」的 LoRA 外掛,名字都叫 Detailer 或 Enhancer,到底差在哪?我用同一批題目、同一組亂數種子,把三款外掛在三個 Krea2 模型上各跑一輪——結果它們根本是三種不同的職業。
AI Image我把同一套 35 組 prompt 交給 Claude,讓它指揮 Codex(ChatGPT 訂閱)、Antigravity(Gemini 訂閱)、Grok 三個 CLI 各生一張圖。我先盲測選出每組我主觀覺得好看的一張,Claude 再用「有沒有照 prompt 做」評一輪——兩份結果一對照,發現「聽話」和「好看」真的不是同一件事。
AI ImageBoogu、Z-Image、Qwen 2512、Ideogram 4 用同一套 35 組 prompt 同 seed 對決(Krea2 Turbo 壓陣)。結果評比做到一半變成偵探劇:Ideogram 4 拒繪率 17%,被擋的全是無害題——追到最後,兇手不是內容審查。
AI Image用同一套 35 組 prompt、同一組 seed,實測 Krea2 官方 Turbo 與四個社群微調(Kreamania v3/v4/v5、Muse)。誰最有藝術感?誰最適合商業廣告?附完整對照圖與 prompt 分享。
Dev NotesSame handwritten logo, same brief, three AI agents — and four different font families came back. What the experiment revealed about the ceiling of AI type design, why picking a base font is a taste decision, and why licensing is the real deep water. 同一個手寫 Logo 交給 Claude、Codex、Grok 三個 AI 做品牌字型, 最後學到的是:字型設計裡哪些能自動化、哪些不能。
Behind the Work我描述了一隻圍著紅圍巾、在極光下奔跑的企鵝,AI 把它做成了一個零素材的偽 3D 跑酷遊戲,還寫了 51 項自動化測試給自己打分數——全數通過。然後我玩了三分鐘:太難了,企鵝還會原地轉圈。這篇是那之後發生的事:三個數學錯誤、一隻用人類反應速度玩遊戲的機器人,和一個「輸了也在前進」的魚幣經濟。
AI Image過去我的每支 AI 影片都是自己一步步做的:想分鏡、寫 prompt、生圖、餵影片模型、下載、剪接。這次換個做法——我只說了一句「我想要奇幻花朵開花、每段不同的花、連續變幻、4:5」,剩下交給 AI agent 指揮四個工具接力完成。這篇記錄這場「完成度測試」的結果:哪些環節 agent 自己搞定了,哪些地方還是得有人類在場。
Dev NotesA seven-day trial, a new model, and one question worth asking on day one: what's actually in your hands? The model list won't tell you.
Behind the WorkI turned my retro tank game into a hospital infection-control game — where using the wrong tool isn't just weak, it's the lesson. This time the reused engine made it fast, 45 automated checks went green, and yet the scariest bug wasn't in the game at all: for a while my tests were quietly grading a different game entirely.
Behind the WorkI can't read code. But I loved tank battle games on the Famicom as a kid, so I described the game I remembered in plain language and let AI build, test, and deploy it. All 43 automated checks passed — then ten minutes of actually playing found the bug none of them caught.
Behind the WorkI wrote a puzzle-game prompt for my AI workshop students, then had to prove it actually works. OpenAI's new model built the whole game in a day — four interlocking puzzles, 32 automated tests, GPT Image 2 scene art — and burned nearly my entire weekly quota doing it. It died one step before deployment. So a second AI picked up where the first one stopped, reading its predecessor's work notes from a shared memory. Those notes had one problem: every Chinese character had turned into "????".
AI Image用同一個 prompt 比模型?那不公平。這次讓 Midjourney 和 Krea 2 各自用最擅長的方式出圖,11 組併排對比,畫出一張「什麼時候該用誰」的使用場景地圖。
AI ImageERNIE-Image Turbo、Qwen-Image-2512、Boogu Base、Z-Image、KREA 2 Turbo — 用 53 組 prompt 全面測試繁體中文文字渲染與美感表現。「濱」這個字,有模型寫得出來嗎?
Behind the Work一個蒸汽龐克愛好者的機械變形夢——從 SDXL + LoRA 逐幀生圖到 Kling 動態生成,最花時間的竟然是金屬碰撞的配音。
Behind the Work一個敦煌舞愛好者遇見「敦煌萌娃」LoRA——小花神的故事就這樣在腦中完整浮現。Flux 逐幀生圖 × Kling 動態 × 剪映剪輯的早期 AI 影片製作紀錄。
Behind the Work七扇門、七個不該存在的世界——用 ComfyUI Flux 逐幀生成的超現實奇幻短片製作全紀錄。
AI ImageKREA 2 官方大更新——原本 4 個 LoRA 換成 9 個,風格從復古動漫、雨窗、塔羅牌到兒童塗鴉都有。同一個 Prompt 跑 10 種風格,看看哪個最對你的胃口。
AI Image用 ComfyUI 實測 Boogu Image 0.1 三個變體(Base、Turbo、Edit),分享 10 組 prompt 與生成結果,誠實比較文字渲染、攝影品質、風格轉換的表現與限制。
Behind the WorkHow a screenshot became a six-continent AI film — reverse-engineering prompts with Gemini, generating with MidJourney, and animating with Seedance (即夢).
AI ImageKREA 2 是目前美學最強的開源生圖模型。實測中文 Prompt 能不能直接用?四種官方 LoRA 風格差在哪?附完整 Prompt 和 ComfyUI 簡易安裝說明,讓你直接跑。
Tool ShowcaseHow I turned a 920-line HTML page of MidJourney prompts into a full prompt planning tool — with an interactive map, style templates, route planner, and 97 landmarks across six continents.
Behind the WorkHow a single Midjourney prompt of a cat in a metallic tunnel became an 8-episode surreal short film series — each world built from one visual idea and a lot of re-rolls.
Behind the WorkMid-Autumn moonlight, red turtle cakes for blessings, and blueberry tanghulu under the stars — the second half of the miniature kitchen series, where every pastry fought back.
Behind the WorkHow a love for street food became a miniature cooking series — tiny hanfu figures chasing chickens, stirring woks, and pulling noodles in a world the size of a teacup.
Behind the WorkA Kling Challenge entry became a whole parallel universe — where 2026's tech runs on brass gears, steam projectors, and ornithopters over Gothic London.
Behind the WorkTwo jazz parades, one city, zero humans allowed — how the Cat Parade Series traded narrative for pure rhythm.
Tool ShowcaseCompose motifs into seamless tiles, inspect with seam heatmap, and export print-ready PNG/PDF. All in the browser.
Dev NotesI found a repo that turns OpenStreetMap data into 3D buildings. So I put it on my site — with search, fog, and Morandi colors.
Dev NotesI saw a dragon follow a cursor across a medieval page, and the text just moved out of the way. So I made my own version — with magnolia flowers.
Tutorial從基礎到進階,教你用 GPT-image 生成超可愛迷你分身。場景道具、動作設計、手寫塗鴉、微縮模型⋯⋯26 種變化一次拆解,附完整 Prompt 範例。
Tutorial用 GPT-image 和 AI 生圖工具,一張照片就能變出 10 種風格!迷你分身、漢服寫真、手繪插畫、微縮世界⋯⋯附完整 Prompt 範例與拆解教學。
Behind the WorkHow a fairytale about glowing seeds became a short film about green energy, renewal, and the quiet power of believing the earth still wants to grow.
Behind the WorkA Gothic Halloween waltz where every character has a story that doesn't end well — and that's exactly the point.
Behind the WorkHow the Cat Duke's second chapter became a Christmas waltz about composure, longing, and the courage to extend a hand.
Behind the WorkA neo-punk K-pop music video about a girl in a neon maze — how illusion, control, and self-awakening became a song.
Tool Showcase19 tools in one place — from AI image generation to print-ready sticker sheets. Built because I got tired of switching between 10 different apps.
Tool ShowcaseUpload a photo, describe what you want to wear, and see the result in seconds. No changing room needed.
Behind the WorkA short film born from clinical reality — how I turned the invisible threat of antimicrobial resistance into an AI-driven cinematic experience.
Behind the WorkHow a foggy steampunk city and a cat with a brass cane became a waltz — the story behind The Duke of Amber Lights.