Tap, and just say it
Image, video, audio and text nodes, storyboard scripts, prompt nodes, and the XinYu chat on the right — all of them now have a microphone next to the send button.
Tap once to start, tap again to stop. What you say lands straight in the box, appended to whatever you had already written.
Nothing is sent automatically. Look it over, fix a word, add your @ references, and send when you're ready. For long prompts, say a chunk, pause, then keep going.
Chinese and English
To the left of the microphone there's a 中 / EN toggle. Pick one and it recognizes in that language; your choice is remembered for next time.
Speaking a Chinese prompt with English model names or craft terms mixed in? Stay on 中 and say it naturally — common creative terminology is boosted.
No credits
Recognition happens in your browser. Your audio never reaches our servers, and it costs no credits.
Chrome or Safari work best. The first time you tap the microphone your browser will ask for permission — allow it. Safari needs Dictation or Siri enabled in system settings. If your browser can't do it, the button simply won't appear, so you'll never tap something that does nothing.