← All Categories

Text-to-Speech

172 tools

Image to AI voice

Convert images to text and AI voice.

This website converts images to text and AI voice. It allows users to upload an image and extract the text content from it.

FreemiumFree tier

Swiftink

Contract Signing Made Simple

Swiftink is an advanced AI platform that transforms audio and video content into precise text transcriptions. It offers best-in-class ASR via OpenAI Whisper, domain-aware capabilities, and supports over 95 languages. Swiftink provides speech-to-text, dictation, and voice recognition services through its API and plugin.

Contact

AnyToSpeech

Convert Any Text to Speech Instantly

AI Text to Speech Converter A clean and simple AI text-to-speech solution. An easy way to convert text, pdf, docs, scan, image to speech. Features TEXT TO SPEECH BLOG TO PODCAST PDF TO SPEECH SCAN or IMAGE TO SPEECH URL TO SPEECH Text Input Options Text Document URL Image Voice Options English (US) Voices: Nova, Onyx, Shimmer, Fable, Echo, Alloy, Erica, Emma, Sophia, Charlotte, Amelia, Evelyn, Grace, Clara, David, Jack, Harry, Richard, Albert, Henry, William, Daniel, Oliver. English (UK) Voices: Jacob, Sebastian, Mateo, Samuel, Joseph, Olivia, Amelia, Isla, Lily, Freya, Daisy, Sienna. English (India) Voices: Krishna, Aarav, Dhruv, Arjun, Maya, Lakshmi, Jaya, Parvati. English (Australia) Voices: Adam, Ashton, Nathan, James, Harvey, Xavier, Zoe, Bella, Hannah, Penelope, Luna, Evie. Afrikaans (South Africa) Voice: Amahle. Arabic Voices: Amir, Hassan, Omar, Abdul, Fatima, Aisha, Inaya, Salma. Other Voices: I...

Free

kardome.com

Voice AI that hears and understands like people do

Kardome’s voice user interface technology clusters speech signals based on location, giving clear real-time voice command input and audio output in any environment. Kardome’s AI technology offers an all-in-one solution for manufacturers and OEMs looking to improve their existing speech recognition systems. Kardome’s break through technology improves voice recognition accuracy in challenging soundscapes, transforming voice UI from a cloud-dependent experience to a secure, real-time, and customizable user experience driven by neural network technology that is deployable to any smart device.

Contact

ClearCypherAI

ClearCypher LLC is a company that builds Generative AI products, including Audio to Audio (T2T) speech engine, Text to Audio (T2A) speech engine, and Audio to Text (A2T) transcription engine. They offer machine learning solutions specializing in automatic speech recognition, machine translation, optical character recognition, and speaker identification. Their platform provides language technology solutions for processing audio, video, image, and text content, delivering enterprise-grade language translation and voice biometrics.

FreemiumFree tier

Awesome-Chinese-LLM

整理开源的中文大语言模型,以规模较小、可私有化部署、训练成本较低的模型为主,包括底座模型,垂直领域微调及应用,数据集与教程等。

整理开源的中文大语言模型,以规模较小、可私有化部署、训练成本较低的模型为主,包括底座模型,垂直领域微调及应用,数据集与教程等。

FreeFree tier

Audiosonic

Transform Text into Realistic Audio with Audiosonic by Writesonic

Audiosonic is an AI-powered text-to-speech tool developed by Writesonic that transforms written text into realistic, human-like audio 1. Its core purpose is to provide high-quality, engaging audio content quickly and easily, eliminating the need for expensive voice actors and recording studios 12. Key features include: Realistic, human-like audio generation using advanced deep learning algorithms 12 Support for over 30 languages and dialects 18 Customizable voice settings (gender, accent, tone, speed, pitch) 112 Instant AI voice generation 12 Commercial use clearance for generated audio 12 Potential applications include marketing and advertising, education, podcast production, accessibility solutions, software demos, and content repurposing. Audiosonic's unique selling points are its high-quality natural-sounding audio, extensive multilingual support, ease of use, instant audio generation, and seamless integration with Writesonic 112. Technically, Audiosonic is a cloud-based SaaS application requiring an internet connection 15. It integrates fully within the Writesonic platform, streamlining the content creation process 112. While specific awards or recognition are not documented, Audiosonic was released in September 2023 and continues to be improved 612. Its advanced capabilities and integration with Writesonic position it as a powerful tool for businesses and content creators seeking to efficiently produce high-quality audio content.

FreemiumFree tier

ailearning

AiLearning:数据分析+机器学习实战+线性代数+PyTorch+NLTK+TF2

AiLearning:数据分析+机器学习实战+线性代数+PyTorch+NLTK+TF2

FreeFree tier

AudioBook Bot

Transform Your Text Into Audiobooks Effortlessly with AudioBook Bot!

AudioBook Bot is an AI-powered platform designed to convert written text into high-quality audiobooks 12. It offers a fast, affordable, and user-friendly solution for audiobook creation, reducing the time and cost typically associated with traditional methods 12. The tool's primary function is the automated generation of audiobooks from text input using generative AI for text-to-speech conversion 124. It supports multiple voices and characterizations, creating engaging listening experiences with dynamic narrations 12. Key features include: Text-to-speech AI: Converts text into natural-sounding speech using advanced AI 12[4](https://www.aibase.com/tool/30333]. Multi-voice support: Offers a selection of over 120 licensed voices for character-rich narrations 2. Customizable settings: Allows users to customize aspects of the audiobook generation process 12. Easy uploading: Provides a straightforward process for uploading written work 1. Potential use cases span various fields: Self-publishing authors: Enables authors to produce their own audiobooks 1. Educational resources: Facilitates the creation of audio versions of educational materials 1. Podcasts and storytelling: Supports the production of podcasts and audiobooks for storytelling 1. Marketing and promotions: Can be used to create audio marketing materials 1. Content repurposing: Allows existing written content to be repurposed into audio format 1. AudioBook Bot's advantages include its ease of use, speed, and affordability compared to traditional audiobook creation 12. The AI-driven process reduces production time and cost 1. The availability of multiple voices enhances the listening experience 12. Regarding technical specifications, the pricing model is based on character count, with a standard package costing $100 per 100,000 characters 2. Information on integration capabilities, achievements, awards, and recent updates is limited. The platform was added to Creati.ai on June 6, 2024 1 and to WhatTheAI on May 10, 2024 2.

Contact

reccloud.cn

新一代AI音视频处理平台 | Next-gen AI audio/video processing platform

RecCloud is a leading AI audio and video processing platform that offers a range of tools for content creation and editing. It includes features like AI speech-to-text, AI subtitles, AI text-to-speech, and AI video translation. The platform is designed to be user-friendly and accessible online.

FreemiumFree tier

Inpodcast AI

Revolutionize Your Content with AI-Powered Podcasts

Inpodcast AI is a podcast creation suite designed to simplify and expedite the podcast production process, transforming written content into high-quality audio podcasts using AI-powered text-to-speech (TTS) technology and intelligent audio processing 127. It aims to make professional-level podcasting accessible to a broader audience, regardless of technical expertise or access to professional recording equipment 12. Key features and capabilities include: Document to Podcast: Converts documents in formats like PDF, Docx, Markdown, and TXT into podcasts 127, with AI voice synthesis 127, multilingual processing supporting over 70 languages 12, and customizable scripts 1. Script to Podcast: Transforms podcast scripts directly into professional audio 1, incorporating smart pacing and segmentation 1, providing access to a library of over 100 unique voices 12, and offering integrated sound effects and background music 1. Text to Speech: Converts text input into natural-sounding speech with premium audio quality 12, supporting over 30 languages and easy content import 1. Potential use cases and applications span across: Education and Training: Converting lecture notes, language learning materials, and teaching outlines into audio 1. Corporate Communications: Creating internal news podcasts, audio training courses, and product introduction audio 1. Personal Creation: Enabling bloggers, authors, and podcast enthusiasts to easily produce podcasts from articles, books, and ideas 1. Inpodcast AI's unique selling points include multi-format support for documents (PDF, Docx, Markdown, TXT) 12, an extensive voice library (over 100 voices) 12 12, and a freemium pricing model 12. The provided text does not detail specific technical requirements or integration capabilities. The provided sources mention user reviews on a third-party site 8, but don't provide specifics on awards, industry recognition, or the content of those reviews, and no specific details about recent developments are offered.

Contact

best-of-ml-python

🏆 A ranked list of awesome machine learning Python libraries. Updated weekly.

🏆 A ranked list of awesome machine learning Python libraries. Updated weekly.

FreeFree tier

RAG_Techniques

This repository showcases various advanced techniques for Retrieval-Augmented Generation (RAG) systems. Each technique has a detailed notebook tutorial.

This repository showcases various advanced techniques for Retrieval-Augmented Generation (RAG) systems. Each technique has a detailed notebook tutorial.

FreeFree tier

LazyTyper

Voice typing powered by 12 AI models, 3x faster, always free.

LazyTyper is a free, super-fast, and highly accurate voice typing application powered by Whisper and other advanced AI speech models. It offers 12 professional speech models, including 5 fully local (on-device) options, enabling users to convert speech to text 3 times faster than manual typing with 90% accuracy. The app supports multilingual dictation, handles accents and technical terms, and is designed to be lightweight, working efficiently on Windows, macOS, and Linux. It is completely free, without ads, and prioritizes user privacy by sending voice data directly to chosen API providers without storing it on LazyTyper's servers.

FreeFree tier

Loquendo

Generate realistic Spanish voiceovers online—fast, simple, and affordable.

GeneradorDeVoz.com is an AI-powered Spanish text-to-speech (TTS) platform that converts written text into natural, human‑sounding audio. It offers realistic voices across Spanish accents and dialects (Spain, Mexico, Argentina, Colombia, and more), with granular controls for speed, pitch, and pauses via SSML. Fully browser-based, it lets you paste text, choose a voice, customize, and download MP3/WAV/OGG in seconds. A free tier covers basic needs, while affordable premium plans enable longer scripts, commercial rights, higher fidelity (up to 48kHz and 320kbps), and watermark-free downloads—ideal for creators, educators, and businesses.

Contact

Claudet: Claude.ai Voice Input

Voice input for Claude.ai

Claude.ai is an AI assistant designed to be helpful, harmless, and honest. It can assist with a variety of tasks, including summarizing text, answering questions, generating creative content, and providing helpful recommendations. It is developed by Anthropic and focuses on safety and ethical AI development.

Contact

DictationDaddy - Speak To Type

AI Speech to Text That Gets It Right the First Time

Dictation Daddy works anywhere across all the apps. You can just speak and it will create a 100% accurate transcript.

Contact

BettaFish

微舆:人人可用的多Agent舆情分析助手,打破信息茧房,还原舆情原貌,预测未来走向,辅助决策!从0实现,不依赖任何框架。

微舆:人人可用的多Agent舆情分析助手,打破信息茧房,还原舆情原貌,预测未来走向,辅助决策!从0实现,不依赖任何框架。

FreeFree tier

Voice Inbox

Inbox Reimagined: Your go-to for jotting down thoughts on the go.

Voice Inbox is a tool designed for quickly capturing thoughts on the go. It transcribes spoken words with human-level accuracy and saves them to a journal, allowing users to focus on expressing themselves and managing tasks. It integrates with Obsidian for seamless note-taking.

FreemiumFree tier

Coachchat

AI voice tutor for personalized coaching anytime, anywhere

Coachchat is an AI voice tutor platform that provides personalized coaching on any topic. It enhances the learning experience with AI voice interaction, offering personalized lessons and guidance accessible 24/7 from anywhere in the world. It helps improve skills and overcome challenges through chat-based coaching.

FreemiumFree tier

Sayline

Stop typing. Just say the line and watch it appear.

Sayline is a native macOS application designed for private, local voice dictation in any text field. It allows users to replace manual typing with voice commands using global hotkeys across various applications like Gmail, Slack, VS Code, or Notes. Utilizing on-device processing technologies (NVIDIA Parakeet and MLX), Sayline ensures uncompromised security and privacy by keeping all audio and data local to the user's Mac, never sending it to the cloud. Sayline is engineered to boost productivity, claiming to be 4x faster than manual typing.

FreemiumFree tier

datasets

🤗 The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data manipulation tools

🤗 The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data manipulation tools

FreeFree tier

langextract

A Python library for extracting structured information from unstructured text using LLMs with precise source grounding and interactive visualization.

A Python library for extracting structured information from unstructured text using LLMs with precise source grounding and interactive visualization.

FreeFree tier

Lovevoice

Turn text into natural-sounding speech with 300 voices in 70+ languages—no subscription required.

Lovevoice AI is a powerful text-to-speech (TTS) and AI voice generator that transforms written text into natural-sounding audio with nearly 300 realistic voices across 70+ languages. Built for creators and businesses, it captures context, tone, and emotion to deliver lifelike speech for YouTube, TikTok, podcasts, ads, training, and more. With customizable voice controls, fast processing for long-form content, MP3 downloads, and commercial rights on paid plans, Lovevoice AI helps teams scale content while maintaining a consistent brand voice. A one-time purchase credit system—no subscriptions—keeps costs predictable and simple.

Contact
PreviousPage 5 of 8Next