Voice · Tested
Eleven v4 Voiceover for a Product Video: What Fits in 15 Seconds

A voiceover turns a slideshow of product photos into a small advert. The hard part is not the voice. It is time: a 15-second video with three shots gives each line about five seconds, and a sentence that reads short on paper can take seven seconds to say.
ElevenLabs released Eleven v4 on September 28, 2026, and Ciyo's agent now uses it for narration. On September 29, 2026 we made a 15-second launch video for Low Tide, a mug from Tidewell, a fictional pottery studio, and measured where every word landed. The voice was clear. The timing needed two rounds.
What Eleven v4 is
Eleven v4 is ElevenLabs' newest text-to-speech model, released with a faster Eleven v4 Turbo for real-time agents. ElevenLabs says both support more than 90 languages and that you can direct the delivery with inline tags. In Ciyo, the agent's Narration tool uses Eleven v4 with a fixed list of ten voices; Sarah, a warm and calm female voice, is the default.
In Ciyo the narration is not a separate file on your board. The agent writes the script by chapter, generates one clip per chapter, and mixes them into a finished video that it saves on the board.
Posted 2026-09-28. ElevenLabs' launch post introduces Eleven v4 and Eleven v4 Turbo as its “fastest and most emotive voice models yet” and says they rank first with Artificial Analysis. We did not check that ranking; our test below is about timing a voiceover for a short product video.
Start with photos that match
We asked for three 9:16 photos of the same mug: alone on a shelf, lifted with coffee, and three in a row. The agent made the first photo, then used it as the reference for the other two so the glaze and the shape stay the same. The three cost 8 credits.

I run Tidewell, a small pottery studio. We launch a new mug called Low Tide on Saturday. Please make three 9:16 product photos of it for a short vertical video: 1) the mug alone on a driftwood shelf by a window, sea-glass green glaze over speckled cream clay, soft morning light; 2) a hand lifting the same mug full of coffee; 3) three of the same mugs in a row on the shelf. Keep the mug identical in all three. Natural colours, no text, no logos.Write the script by chapter
We gave the agent the chapters, the voice and the exact words. It generated one narration clip per chapter and measured each before it assembled the video. Then it told us about the problem: “Chapter 1's narration came out at 6.53 s, longer than its 5 s slot.”
It did not speed the voice up or cut words without asking. It let the first line run 1.5 seconds into the second photo, moved the later lines along, and reported the timing. The three clips cost 1 credit each.

The three photos match. Now make a 15-second vertical video (1080x1920) from them with an English voiceover in the voice Sarah. Chapters: 0 to 5 s photo 1, 5 to 10 s photo 2, 10 to 15 s photo 3, each with a slow push-in. Voiceover, word for word: "Meet Low Tide. Sea-glass glaze over speckled clay, thrown by hand at Tidewell. It holds your morning coffee and feels like the shore. Low Tide mugs arrive Saturday at tidewell.shop." No music. Over the last 3 seconds show the text "Low Tide - Saturday - tidewell.shop". Tell me how long each narration clip is and how many credits the video used.Where the voice actually sat
We measured the finished video with ffmpeg. The voice ran from 0.26 to 6.49 seconds, 6.78 to 10.10 and 10.97 to 14.54. So the first line talks over the cut at 5 seconds, and the second line ends right on the cut at 10.
The numbers give a simple rule. The first line had 13 words and took 6.5 seconds with its pauses, about 2 words per second. The second had 10 words in 3.3 seconds, because it had no full stop in the middle; the full stop after “Meet Low Tide.” alone added a 0.44-second pause. Plan about 8 to 10 words for a 5-second shot, and count each full stop or colon as roughly half a second.

Fix one chapter, not the whole script
We shortened the first line and asked the agent to remake only that clip. Even the shorter line came out a little long, so the agent trimmed the pause after “Meet Low Tide:” and sped the clip up by 3%, to 4.62 seconds. That cost 1 credit, and the other two clips were reused.
In the second version every line sits inside its own photo: 0.33 to 4.86 seconds, 5.38 to 8.70 and 10.37 to 13.94. The loudness is −16.5 LUFS with peaks at −1.3 dB, close to the −16 LUFS that the agent's video recipe aims for.
Two fixes, please. 1) Keep each line inside its own photo. Replace only clip 1 with this shorter line and make just that clip again: "Meet Low Tide: sea-glass glaze, thrown by hand at Tidewell." Reuse clips 2 and 3 as they are, and start each clip about 0.3 s after its photo appears. 2) The end text is small and sits at the top, where the Instagram header covers it. Make it large and centred at about 60% of the frame height, on two lines: "Low Tide mugs" and "Saturday - tidewell.shop", with a soft dark band behind it. Keep the video exactly 15 s and tell me the new clip length and the credits used.Check the end card on a phone-sized frame
The first end text was a small grey band near the top, between about 220 and 325 pixels of 1,920, where the app's header can cover it. Our fix made it large and centred at 60% of the height, which put it on top of the three mugs. A third request moved it onto the shelf, at 72%, above the bottom 20% where Instagram places the caption. That change cost nothing, and the audio stayed identical.

The voice timing is right now. One last change: at 60% the end text sits on top of the three mugs and hides the product. Move both lines down onto the wooden shelf below the mugs, at about 72% of the frame height, keep them clear of the bottom 20% where Instagram puts the caption, and make the dark band behind the text a bit stronger. Keep everything else the same and save it as v3.What the whole video cost
Twelve credits in total: 8 for the photos, 3 for the first narration and 1 for the remade chapter. Building the video from photos, and every later change to the text and timing, cost nothing, because no new video was generated.
| Check | Version 1 | Version 3 (final) |
|---|---|---|
| Length | 15.000 s | 15.000 s |
| Every line inside its photo | No (first line runs to 6.49 s) | Yes |
| Loudness | −16.6 LUFS | −16.5 LUFS |
| End text | Small, at the top | Large, on the shelf, above the caption area |
| Credits | 11 (8 photos + 3 narration) | 12 in total |

Eleven v4 voiceovers
What is Eleven v4?
ElevenLabs' text-to-speech model released on September 28, 2026, with a faster Turbo version. ElevenLabs says it supports more than 90 languages and inline delivery tags.
How many words fit in a 5-second shot?
In our test the voice Sarah said about 2 words per second in a line with a pause and 3 in a line without one. Plan about 8 to 10 words for 5 seconds, and count each full stop or colon as roughly half a second.
Can I choose the voice in Ciyo?
Yes. Name one of the Ciyo voices in your request: Sarah, Alice, Matilda, Jessica, Lily, George, Daniel, Brian, Eric or Will. You can also write a hard name's pronunciation in IPA.
Do I get a separate audio file?
No. In Ciyo the narration is mixed into the finished video, which the agent saves on your board.
Make a narrated product video in Ciyo
Give the Ciyo Agent your photos, chapters and exact words, then check the timing before you post.