Main takeaways
Images = grids of numbers: AI spots patterns, from edges to objects
Sound = pictures of vibrations: turn audio into spectrograms, then use image AI
The big insight: text, images, and audio all become embeddings
Multimodal AI: ChatGPT, Claude, and Gemini see images because everything speaks the same mathematical “language”
Hands-on: you trained your own AI with Teachable Machine
Critical thinking: should AI voices and images need watermarks, disclosure or consent?






















