Voice · Tested

Eleven v4 Voiceover for a Product Video: What Fits in 15 Seconds

An ivory ceramic cone lying on its side with three thin rings rising from its opening, the last one amber, beside a graphite block
Original abstract editorial artwork generated in Ciyo with GPT Image 2.5 Sunburst.

A voiceover turns a slideshow of product photos into a small advert. The hard part is not the voice. It is time: a 15-second video with three shots gives each line about five seconds, and a sentence that reads short on paper can take seven seconds to say.

ElevenLabs released Eleven v4 on September 28, 2026, and Ciyo's agent now uses it for narration. On September 29, 2026 we made a 15-second launch video for Low Tide, a mug from Tidewell, a fictional pottery studio, and measured where every word landed. The voice was clear. The timing needed two rounds.

What Eleven v4 is

Eleven v4 is ElevenLabs' newest text-to-speech model, released with a faster Eleven v4 Turbo for real-time agents. ElevenLabs says both support more than 90 languages and that you can direct the delivery with inline tags. In Ciyo, the agent's Narration tool uses Eleven v4 with a fixed list of ten voices; Sarah, a warm and calm female voice, is the default.

In Ciyo the narration is not a separate file on your board. The agent writes the script by chapter, generates one clip per chapter, and mixes them into a finished video that it saves on the board.

Start with photos that match

We asked for three 9:16 photos of the same mug: alone on a shelf, lifted with coffee, and three in a row. The agent made the first photo, then used it as the reference for the other two so the glaze and the shape stay the same. The three cost 8 credits.

The Ciyo canvas with three vertical photos of a sea-glass green mug on a driftwood shelf by a window: alone, lifted by a hand, and three in a row
The three product photos for the video. September 29, 2026.
The first message to the Ciyo Agent
I run Tidewell, a small pottery studio. We launch a new mug called Low Tide on Saturday. Please make three 9:16 product photos of it for a short vertical video: 1) the mug alone on a driftwood shelf by a window, sea-glass green glaze over speckled cream clay, soft morning light; 2) a hand lifting the same mug full of coffee; 3) three of the same mugs in a row on the shelf. Keep the mug identical in all three. Natural colours, no text, no logos.

Write the script by chapter

We gave the agent the chapters, the voice and the exact words. It generated one narration clip per chapter and measured each before it assembled the video. Then it told us about the problem: “Chapter 1's narration came out at 6.53 s, longer than its 5 s slot.”

It did not speed the voice up or cut words without asking. It let the first line run 1.5 seconds into the second photo, moved the later lines along, and reported the timing. The three clips cost 1 credit each.

The Ciyo agent panel with three narration steps and a note that chapter 1 came out at 6.53 seconds, longer than its 5-second slot
The agent measures each narration clip before it builds the video.
The video request
The three photos match. Now make a 15-second vertical video (1080x1920) from them with an English voiceover in the voice Sarah. Chapters: 0 to 5 s photo 1, 5 to 10 s photo 2, 10 to 15 s photo 3, each with a slow push-in. Voiceover, word for word: "Meet Low Tide. Sea-glass glaze over speckled clay, thrown by hand at Tidewell. It holds your morning coffee and feels like the shore. Low Tide mugs arrive Saturday at tidewell.shop." No music. Over the last 3 seconds show the text "Low Tide - Saturday - tidewell.shop". Tell me how long each narration clip is and how many credits the video used.

Where the voice actually sat

We measured the finished video with ffmpeg. The voice ran from 0.26 to 6.49 seconds, 6.78 to 10.10 and 10.97 to 14.54. So the first line talks over the cut at 5 seconds, and the second line ends right on the cut at 10.

The numbers give a simple rule. The first line had 13 words and took 6.5 seconds with its pauses, about 2 words per second. The second had 10 words in 3.3 seconds, because it had no full stop in the middle; the full stop after “Meet Low Tide.” alone added a 0.44-second pause. Plan about 8 to 10 words for a 5-second shot, and count each full stop or colon as roughly half a second.

Timeline of speech in two versions: in version 1 the first line crosses the 5-second cut; in version 2 all three lines sit between the cuts
Speech in the two versions, against the photo cuts at 5 and 10 seconds. Measured by us.

Fix one chapter, not the whole script

We shortened the first line and asked the agent to remake only that clip. Even the shorter line came out a little long, so the agent trimmed the pause after “Meet Low Tide:” and sped the clip up by 3%, to 4.62 seconds. That cost 1 credit, and the other two clips were reused.

In the second version every line sits inside its own photo: 0.33 to 4.86 seconds, 5.38 to 8.70 and 10.37 to 13.94. The loudness is −16.5 LUFS with peaks at −1.3 dB, close to the −16 LUFS that the agent's video recipe aims for.

Our fix request
Two fixes, please. 1) Keep each line inside its own photo. Replace only clip 1 with this shorter line and make just that clip again: "Meet Low Tide: sea-glass glaze, thrown by hand at Tidewell." Reuse clips 2 and 3 as they are, and start each clip about 0.3 s after its photo appears. 2) The end text is small and sits at the top, where the Instagram header covers it. Make it large and centred at about 60% of the frame height, on two lines: "Low Tide mugs" and "Saturday - tidewell.shop", with a soft dark band behind it. Keep the video exactly 15 s and tell me the new clip length and the credits used.

Check the end card on a phone-sized frame

The first end text was a small grey band near the top, between about 220 and 325 pixels of 1,920, where the app's header can cover it. Our fix made it large and centred at 60% of the height, which put it on top of the three mugs. A third request moved it onto the shelf, at 72%, above the bottom 20% where Instagram places the caption. That change cost nothing, and the audio stayed identical.

Three end frames of the mug video side by side: small text at the top, large text over the mugs, and large text on the shelf below the mugs
The last frame of versions 1, 2 and 3. Ciyo Agent output, arranged by us.
The last change
The voice timing is right now. One last change: at 60% the end text sits on top of the three mugs and hides the product. Move both lines down onto the wooden shelf below the mugs, at about 72% of the frame height, keep them clear of the bottom 20% where Instagram puts the caption, and make the dark band behind the text a bit stronger. Keep everything else the same and save it as v3.

What the whole video cost

Twelve credits in total: 8 for the photos, 3 for the first narration and 1 for the remade chapter. Building the video from photos, and every later change to the text and timing, cost nothing, because no new video was generated.

Voiceover checks, first version and final version
CheckVersion 1Version 3 (final)
Length15.000 s15.000 s
Every line inside its photoNo (first line runs to 6.49 s)Yes
Loudness−16.6 LUFS−16.5 LUFS
End textSmall, at the topLarge, on the shelf, above the caption area
Credits11 (8 photos + 3 narration)12 in total
The Ciyo agent panel describing version 3: text at 72% of the height, a darker band, no credits used
The agent's report on the final version.

Eleven v4 voiceovers

What is Eleven v4?

ElevenLabs' text-to-speech model released on September 28, 2026, with a faster Turbo version. ElevenLabs says it supports more than 90 languages and inline delivery tags.

How many words fit in a 5-second shot?

In our test the voice Sarah said about 2 words per second in a line with a pause and 3 in a line without one. Plan about 8 to 10 words for 5 seconds, and count each full stop or colon as roughly half a second.

Can I choose the voice in Ciyo?

Yes. Name one of the Ciyo voices in your request: Sarah, Alice, Matilda, Jessica, Lily, George, Daniel, Brian, Eric or Will. You can also write a hard name's pronunciation in IPA.

Do I get a separate audio file?

No. In Ciyo the narration is mixed into the finished video, which the agent saves on your board.

Make a narrated product video in Ciyo

Give the Ciyo Agent your photos, chapters and exact words, then check the timing before you post.