Speech & Audio
The other half of human communication. Hear, understand, speak.
Цена: 9 187 ₽
Длительность: 15 ч
Автор: John Jackson
Программа курса
- Audio Fundamentals — Waveforms, Sampling, Fourier Transform
- Spectrograms, Mel Scale & Audio Features
- Audio Classification — From k-NN on MFCCs to AST and BEATs
- Speech Recognition (ASR) — CTC, RNN-T, Attention
- Whisper — Architecture & Fine-Tuning
- Speaker Recognition & Verification
- Text-to-Speech (TTS) — From Tacotron to F5 and Kokoro
- Voice Cloning & Voice Conversion
- Music Generation — MusicGen, Stable Audio, Suno, and the Licensing Earthquake
- Audio-Language Models — Qwen2.5-Omni, Audio Flamingo, GPT-4o Audio
- Real-Time Audio Processing
- Build a Voice Assistant Pipeline — The Phase 6 Capstone
- Neural Audio Codecs — EnCodec, SNAC, Mimi, DAC and the Semantic-Acoustic Split
- Voice Activity Detection & Turn-Taking — Silero, Cobra, and the Flush Trick
- Streaming Speech-to-Speech — Moshi, Hibiki, and Full-Duplex Dialogue
- Voice Anti-Spoofing & Audio Watermarking — ASVspoof 5, AudioSeal, WaveVerify
- Audio Evaluation — WER, MOS, UTMOS, MMAU, FAD, and the Open Leaderboards
- Итоговое задание