Pronunciation Assessment

by Unknown

Not rated
Website

About

AI-powered English pronunciation scoring at phoneme level. 17MB model, sub-300ms, returns IPA/ARPAbet notation with per-phoneme scores.

Details

Author
Unknown
Categories
Cloud Service, AI, Other

AI-powered English pronunciation scoring at phoneme level. 17MB model, sub-300ms, returns IPA/ARPAbet notation with per-phoneme scores.

Analyze spoken language to provide immediate feedback on pronunciation accuracy and fluency. Identify specific phonetic errors to help learners improve their speaking skills in real-time. Guide users…

# Connect this server (installs CLI if needed) npx -y smithery mcp add fabiosuizu/pronunciation-assessment # Browse available tools npx -y smithery tool list fabiosuizu/pronunciation-assessment # Get full schema for a tool npx -y smithery tool get fabiosuizu/pronunciation-assessment assess_pronunciation # Call a tool npx -y smithery tool call fabiosuizu/pronunciation-assessment assess_pronunciation '{}'

Endpoint:https://pronunciation-assessment--fabiosuizu.run.tools

- apiKey(query) — Your APIM subscription key

- assess_pronunciation— Assess English pronunciation quality from audio.
- check_pronunciation_service— Check if the pronunciation assessment service is healthy and ready.
- get_phoneme_inventory— Get the full phoneme inventory supported by the pronunciation scorer.
- transcribe_audio— Transcribe audio to text with word-level timestamps.
- check_stt_service— Check if the speech-to-text service is healthy and ready.
- synthesize_speech— Generate natural speech audio from English text.
- list_tts_voices— List all available text-to-speech voices with metadata.
- check_tts_service— Check if the text-to-speech service is healthy and ready.
- transcribe_audio_pro— Transcribe audio with Whisper Large V3 Turbo — multilingual STT.
- check_whisper_service— Check if the Whisper STT Pro service is healthy and ready.

# Get full input/output schema for a tool npx -y smithery tool get fabiosuizu/pronunciation-assessment <tool-name>

- pronunciation://scoring-guide— How pronunciation scores work at each granularity level.
- pronunciation://audio-requirements— Supported audio formats, quality requirements, and recording tips.
- whisper://usage-guide— How to use the Whisper STT Pro multilingual transcription service.
- pronunciation://model-info— Performance benchmarks and service characteristics.
- pronunciation://response-schema— Full API response structure with field descriptions.
- pronunciation://example-assessment— Real example showing how to interpret a pronunciation assessment.
- stt://usage-guide— How to use the speech-to-text transcription service.
- tts://usage-guide— How to use the text-to-speech synthesis service.

- analyze_pronunciation(assessment_json) — Analyze pronunciation assessment results and provide actionable feedback.
- create_improvement_plan(assessment_json, learner_level) — Create a structured pronunciation improvement plan based on assessment results.
- compare_attempts(attempt1_json, attempt2_json) — Compare two pronunciation attempts of the same text to track progress.

Turn any language model into a multimodal powerhouse that can generate images, music, videos and more on the fly. Rostro's tools are designed to be used by language models from the ground up, expanding capabilities with minimal context bloat.

MCP-native AI media generation with x402 pay-per-call. Image, video, audio, and music from 6 providers — composable via resource IDs. USDC on Base.

AI audio tools for music producers — stem splitting, vocal removal, BPM/key detection, audio-to-MIDI, format conversion and AI song generation

An MCP server for GPT-SoVITS, providing text-to-speech synthesis, voice cloning, and multi-language support.

A remote MCP server on Cloudflare Workers that generates podcast URLs and rickrolls without authentication, using Cloudflare AI and D1.

An MCP server for the Typecast API, enabling AI-powered voice generation for various content.

Provides speech-to-text, diarization, translation, and text summarization via the Whissle AI API.

An AI voice toolkit with TTS, voice cloning, and video translation, now available as an MCP server for smarter agent integration.

AI Video, Image & Audio Generation with over 150 models

Interact with the Anki flashcard app via the AnkiConnect add-on. Supports audio generation and similarity search.

No reviews yet — be the first

Sign in to leave a review

Use Google, GitHub, or an email account so ratings stay tied to real people.

Email sign in

No reviews posted yet.