How do I sync speech to a face?
7ART Lipsync allows you to turn a portrait or character image into a talking video by synchronizing facial movements with speech or audio.
You can create talking avatars, character videos, and other creative content by combining an image with voice input.
What is Lipsync?
Lipsync is the 7ART tool for creating AI-generated talking videos from images and audio.
You can use it to create:
- Talking portraits
- AI avatars
- Character dialogue videos
- Narrated visual content
- Creative video projects
The final result depends on the selected model, image quality, audio input, and generation settings.
How do I create a Lipsync video?
To create a Lipsync video:
1. Sign in to your 7ART account.
2. Open the Lipsync tool.
3. Add a Portrait Image by uploading an image or selecting one from your library.
4. Add your voice input using one of the available options:
- Voice — use generated audio
- Upload — upload an MP3 or WAV file (maximum 60 seconds)
- Record — record your voice directly (maximum 60 seconds)
6. Choose an available model.
7. Select the desired quality (Higher quality settings may require more credits).
8. Review the credit cost.
9. Click Generate to create your talking video.
5. Add an optional prompt describing the style or mood.
Which Lipsync model should I choose?
7ART provides different models for different types of talking content:
- Kling Avatar — creates a talking avatar from a portrait image.
- HeyGen Avatar IV — creates a talking video from a photo.
- Omnihuman 1.5 — animates a portrait into a talking video.
- Voicengine — re-lipsyncs an existing source video.
Choose the model that best matches your project and desired result.
What affects the result?
The quality of your Lipsync video can be influenced by:
- The clarity of the portrait image
- The quality of the audio input
- The selected AI model
- The quality setting
- The prompt and creative direction
For better results:
- Use a clear portrait where the face is visible.
- Use clean audio with understandable speech.
- Avoid images with faces that are covered or heavily distorted.
- Match the image style and voice style to your project.
What type of image works best?
For the best results, use images with:
- A clearly visible face
- Good lighting
- Front-facing or natural facial positioning
- High image quality
Images with unclear facial details, extreme angles, or heavy editing may produce less predictable results.
What audio works best?
For better speech synchronization:
- Use clear voice recordings.
- Avoid strong background noise.
- Use natural speech pacing.
- Make sure the audio matches the intended character or content.
Uploaded audio should be in MP3 or WAV format and can be up to 60 seconds long.
How much does it cost?
Lipsync generation uses credits.
The credit cost depends on factors such as:
- Selected model
- Quality setting
- Audio length
- Other generation options
Before generating, 7ART shows the required credit cost so you can review it before creating.
What are the limits?
AI-generated talking videos may not always perfectly reproduce every facial movement or expression.
Results can vary depending on:
- Image quality
- Audio quality
- Selected model
- Generation settings
For the best results, use clear inputs and refine your settings when needed.
When should I contact support?
Contact our support team if you experience a technical issue, such as:
- Lipsync generation does not complete
- The tool is not working as expected
- Your image or audio cannot be processed
- You receive an error message
- Your credits are not restored after a failed generation
- Your generated video does not appear in your library
When contacting support, include:
- The email address connected to your 7ART account
- The image and audio input used
- The selected model and quality settings
- A screenshot or error message, if available
This information helps our team investigate the issue faster.
