Skip to content

Bring the narration as an audio file — read aloud by you, or made with a text-to-speech tool of your own. Kutly transcribes what you uploaded and cuts the visuals to it.

In Voiceover the box is a dropzone, headed Drop your narration here, or choose a file over a line reading MP3, M4A or WAV · 8 to 31 min · max 100 MB.

Two of those three are checked before the file uploads: the megabytes and the minutes. Over 100 MB, under 8 minutes and over 31 minutes are each refused with their own reason on screen, and the file is not attached at all.

Neither check is about how the audio was made. A file a text-to-speech tool wrote passes on the same two numbers a recording does.

Nothing checks the format before the upload. The picker also takes AAC, which the hint above does not name. A file that cannot be opened is caught when its length is read instead.

The 8-to-31-minute window is checked a second time once the render starts, so a file can pass on upload and still be turned away. That costs nothing: the render fails and the credits are returned.

Generate — the arrow button at the end of the settings row — is disabled until the upload lands, and the slot beside it says what it is waiting for: Add narration with no file, Finishing upload… while one is on its way.

In Voiceover the settings row still carries Length, and it cannot be changed. There is no Voice setting at all. Your file sets both: the video is as long as the file, and the file is the voice. Kutly synthesises nothing here, which is why there is no voice to pick.

The figure beside Generate is the whole price, with no up to in front of it. The receipt Generate opens totals it under Charged now rather than Held now, because your narration has already fixed the length and there is no difference to come back afterwards. Length and credits has the arithmetic.

The price is worked out from the length, so a file with no readable length has no price to quote.

Some files never give one up. They pass the size check and upload like any other, and where the figure beside Generate would be, the slot reads reading your audio… and stays there.

Generate stays enabled through all of that: what turns it on is the landed upload, not the length. Pressing it opens no receipt. The screen prints “Couldn’t read your audio’s length. Please re-upload or try a different file.” under the composer, and nothing has been charged — the receipt is the only thing here that spends credits, and this refusal comes before it opens.

The way out is a different file. Replace, on the file row, reopens the picker.

The transcript of your recording is what gets cut into pieces. The cutting happens before anything decides what to show. A cut lands on a sentence end or a clause mark — a full stop, a question mark, an exclamation mark, a comma, a semicolon, a colon — or on a pause between words.

Every word ends up in exactly one piece, in the words you spoke.

Read the punctuation you wrote. Those marks come from the transcript, not from the page you read off, so a full stop you typed and did not read never reaches the cut.

Each piece is then walked in order and marked one of two ways: it opens a new visual, or it continues the one before. The decision is made per piece, not per sentence. What shapes the pieces is how you write and read your sentences.

Break a clause you would read for more than twelve seconds. Two bounds hold the pieces: a stretch under two seconds does not become a piece of its own, so a run of short clauses stays together, and a clause that runs past twelve seconds with no mark in it is cut anyway, at an arbitrary point.

Name the thing a figure is about. The footage is searched on a short query written from the piece that opens the visual — the concrete thing, and the setting it is seen in — and exact numbers, dates and sums of money are kept out of it, because they do not image-search. A line that is only a figure leaves that query nothing to name.

Your narration is the whole brief. There is no topic box in this mode: the audio file is the input, and no other text you write reaches the render. Style pack and Music are settings of their own in the row under the box, and Subtitles sits in the second row. Ask for one of them out loud and all you get is a line of narration.

The script holds the words to be spoken and nothing else: every word in the file is heard in the video. A note to the system has nowhere to go here — read aloud, it is narration.

Narration whose whole point is to instruct the system, with no video subject anywhere in it, is refused. The check classifies the transcript of your file before anything is planned or rendered. One adversarial-sounding line inside a long narration is not what it looks for. It looks for a recording with no film in it at all.

A refused narration comes back as a failed render with your credits returned. The video reads Render failed, and where the player would be the screen prints the reason: “We couldn’t use this narration — it reads as instructions for the system rather than the words to be spoken. Upload or paste the narration itself.”