SETUP & SUPPORT
Before your first edit.
CueTwin AI is being prepared for a free-tier Windows beta. Public downloads and paid sales are not open yet. The target is 64-bit Windows 10/11 with 16 GB RAM; CPU operation is supported in the implementation and a compatible GPU is recommended. Clean-machine installation and performance acceptance remain in progress.
Free and Creator
Free includes editing and saving, the bundled QTwin assistant, instant voice cloning, and video export through 1080p. Portrait, square and the built-in ultrawide 1080p presets are included. Creator adds higher-resolution export and professional voice cloning with an active license. The beta starts with the free tier; it does not grant temporary paid access.
The workflow
- Import or record your footage, then save a project.
- Choose CueTwin (built-in) for local editing advice, script polishing and timeline commands. External AI providers remain optional and may charge separately.
- Plan and refine the edit through chat. Review the timeline and preview before exporting.
- Write your narration script. Set up a voice profile using your own clean recording or one you have permission to use.
- Generate narration, listen to it, align it with the footage, and export.
Timeline sections
Right-click a section tab or its block on the timeline to rename, duplicate, or delete it. Duplicate creates an independent copy of its clips and transitions beside the original block. Deleting a nonempty section expands each place it is used into its clips, including transitions inside the section. Use Undo to restore a deleted or duplicated section.
Built-in QTwin assistant
Choose CueTwin (built-in) in the chat selector. Updated builds bundle Qwen3.5 with a local llama.cpp runtime, so chat needs no API key or Ollama installation. CueTwin selects its Lite or Pro model according to available memory and keeps CPU capacity for the editor. Existing provider choices are preserved.
Ask for editing advice, polish English narration, or request timeline changes such as splitting a clip, adjusting pacing or adding titles. Review the result and use Undo when needed. For narration, review the script card before generating speech with a configured voice. Each generated clip retains its own spoken script in Audio metadata and a recovery file beside its audio. After trimming or splitting a take, use Detect script if the saved words no longer match that source region. CueTwin refuses to regenerate a cut from an estimated slice of the original script; review the detected words before generating. Model answers and edits still need review; response speed depends on the PC and task; current CPU-only measurements are not real-time. The bundled model is not yet a validated custom fine-tune, and current local evaluations do not establish perfect English or error-free editing.
Local model troubleshooting
Local (Ollama) remains optional. Updated builds use shorter instructions, a per-request context window capped at 16K, and a smaller tool set for script revision. Ollama already multiplies memory for parallel requests; increasing context can consume GPU memory needed by video and voice. If a request stalls, check the selected model and GPU memory, shorten the conversation, or use the bundled Lite model. Effort and response latency vary by provider and model.
Local voice generation
First-time setup needs Python 3.11 x64 and internet access. Open Free · Setup → Install voice runtime to install the separate speech environments and download the models. The app selects CPU or NVIDIA dependencies and checks the downloaded model files. Clean-machine installation with this complete setup is still being validated.
CPU speech is currently slow: a development-host test took about 54 seconds to synthesize two seconds of speech and used about 11.4 GiB of memory. That host had 128 GB RAM, so this does not certify the 16 GB target. A compatible GPU is recommended; listen to each generated take.
If you already have a configured runtime, open Free · Setup → Locate existing runtime in CueTwin AI and choose its root folder. It must contain the backend and Qwen Python environments. The new location is used the next time the editor opens. Finding those files does not guarantee the models are loaded or your GPU is supported.
Creator activation
Once sales open, your purchase receipt will include a license key. Open Free · Setup, paste the key, and select Activate. Activation checks the license online and stores it securely on Windows. To move devices, deactivate the old device before activating the new one. Activation controls remain unavailable until the store is configured.
A verified license can continue through a temporary connection outage for up to seven days, capped by its expiry. After expiry or revocation, editing, saving, instant cloning and exports through 1080p remain available. Higher-resolution export and professional cloning require an active license. Both desktop export and voice requests enforce these limits.
Record samples for a voice clone
Instant cloning and export through 1080p are free. Professional cloning requires Creator. Open Voice Studio and choose Instant clone or Professional clone. Enter a voice name, choose your microphone, and click Record sample. Speak a complete sentence, then click Stop & add sample. Repeat for as many takes as you need. Each microphone sample has playback controls and a remove button; you can also add existing audio or video files with Choose files.
When your samples are ready, select Create instant preset or Start professional clone. Instant cloning selects a clean reference from your samples. Professional cloning supports English, Chinese, German, Italian, Portuguese, Spanish, Japanese, Korean, French, and Russian. It trains local speaker adapters from at least ten minutes of usable speech after screening and reserving unseen test phrases. CueTwin compares learned-voice and reference-conditioned candidates, then saves only a candidate that passes identity, wording, pacing, clipping, noise, and sentence-ending checks. Record one speaker in a quiet room with a consistent microphone setup. Samples stay on this PC and are not added to the video timeline. If saving fails, use Retry saving to keep the take.
Multilingual cloning and saved samples
For a quick clone, choose Instant clone. It detects the sample language and uses a clear 3–22 second passage as the voice reference. It does not require a ten-minute training recording. In the voice library, Clone saved samples reuses recordings from an earlier attempt, including failed training attempts. Original recordings stay on your PC.
Qwen3 instant and professional voices support English, Chinese, German, Italian, Portuguese, Spanish, Japanese, Korean, French and Russian. The optional local Chatterbox multilingual model adds Hindi, Arabic, Danish, Dutch, Finnish, Greek, Hebrew, Malay, Norwegian, Polish, Swahili, Swedish and Turkish for instant cloning. It requires a separate one-time model download. The saved voice’s Output language selector shows its supported languages. Auto follows the recorded sample’s language; choose a language explicitly for cross-language output.
Instant reference cloning preserves a reusable voice sample but does not fine-tune or quality-certify it. Professional cloning trains an adapter and requires its generated unseen phrases to pass automated quality gates. Listen to generated previews before using narration.
Billing and help
The planned subscription renews automatically. Use Manage subscription in the license panel to open the customer portal. Cancelling stops future renewals; access continues until the paid period ends. Keep your receipt and contact us if activation fails.
For support, include your app version, Windows version, GPU, and the error message. Do not send license keys, payment details, voice samples, or private footage through the contact form unless specifically needed and requested.
Contact Athian Games · Join early access
Smart AI background music
Development builds with the Smart BGM update include a Music tab. Choose the whole video timeline, selected clips, or a custom range, then select Analyse video. Choose Musical direction for melodic music matched to the video, cinematic melody, warm acoustic melody, or ambient atmosphere. A local vision model reads sampled video frames and proposes a lead theme, chord movement, accompaniment and phrase variations, with tempo, key and scene cues. Review or edit the music prompt, select Generate music, listen to the preview, and select Add to timeline.
This workflow needs Ollama with a downloaded vision model, a separate ACE-Step 1.5 installation and music models, and FFmpeg. Use Locate music engine to choose the ACE-Step folder containing its Python environment, then Start engine. After the one-time setup and model downloads, analysis and music generation run locally. Turbo suits fast drafts; SFT and XL options require more compute. Download the selected model before switching it. When other apps occupy your GPU, choose Local engine settings → Music device → CPU. Large SFT scores can take tens of minutes on CPU; CueTwin allows longer CPU renders. Restart the music engine after changing its device.
Generated stereo WAV files and cue plans stay in the project's music folder. Scores up to two minutes render as one continuous performance, with scene changes inside the arrangement. Longer timelines group related scene cues into sections of up to two minutes. The app fits music to the selected duration, adds fades and scene crossfades, and can lower it under narration clips. Sampled-frame analysis does not inspect every frame; narration ducking follows known voice clips rather than detecting speech in every source video. Listen before using a take, and regenerate after changing narration timing. This feature remains part of development builds and has not been certified on a clean Windows installation.
Connect an external AI without desktop control
Current development builds expose CueTwin’s editing tools through a local API and an MCP adapter. A local AI client can inspect the active project, import media, edit clips, use voice tools, save, undo and read progress without taking over your mouse or keyboard. CueTwin briefly holds its own editing controls during each command to prevent conflicting edits; other applications remain usable.
With the updated app running and Node.js 20 or newer installed, add a local MCP server to your AI client. Use node as the command and the absolute path to resources/automation/automation-mcp.cjs inside your CueTwin installation as its argument. For a source checkout, use electron/automation-mcp.cjs. The adapter finds the app automatically through its private local connection file. No extra AI subscription or API key is required for the connection itself.
Ask the assistant to read cuetwin_state and cuetwin_tools first, then run commands with the returned project, sequence and revision context. Poll each job to completion. The app shows the current command and a Request stop button; some operations finish their current step before stopping. App preview captures include only CueTwin. Exports include video, audio, still images, titles, image overlays and supported picture transitions. The app reports unsupported transition arrangements instead of silently replacing them with cuts. Review the exported file before delivery. Cloud-only clients need a local companion to reach the app.
For phrase-by-phrase picture edits, an external assistant can apply a reviewed list of source shots at exact timeline frames in one undoable operation. It can preview the edit plan first, preserve narration within the range, and move all following tracks together when shortening the section. Shot choices still need review against the actual narration and source footage.
Keyboard and resize controls
Updated development builds accept Space from standard keyboards and remapped keypads, including mappings that send a key without a physical scan code. Space controls playback after using a timeline or volume slider; typing fields keep normal text input. Windows multimedia Play/Pause mappings control CueTwin while its window is focused. Map split to Ctrl+B and timeline zoom to Ctrl+Plus or Ctrl+Minus. Hardware profiles still need to target your installed CueTwin executable.
Clip edges have larger trim handles. Trimming sped-up footage preserves its source timing. Panel dividers support dragging and keyboard adjustment: focus a divider and use its arrow keys, holding Shift for one-pixel adjustments. Releasing or cancelling a divider drag restores playback controls.