Edge-deployable AI for zero-shot medical audio classification
This project, conducted as a Research Intern at Samsung PRISM, focused on developing a lightweight 'Student' foundation model for Human Body Sound analysis. The primary objective was to overcome the computational barriers of deploying massive foundation models (like CLAP and CLIP) on edge devices. By implementing a Multi-Teacher Knowledge Distillation framework, the project aimed to compress the 'knowledge' of large multi-modal models into a compact student model (under 50MB) capable of performing zero-shot classification of heart, lung, and bowel sounds on resource-constrained hardware. The project employed a rigorous Data-Centric AI approach, curating over 20,000+ medical audio samples and engineering a novel 3-stage preprocessing pipeline to mathematically isolate bio-acoustic signals from complex environmental noise. This work addresses the critical need for offline, real-time diagnostic tools that can run on low-power devices like digital stethoscopes or wearables, enabling accurate medical assessments in remote or resource-poor settings without requiring cloud connectivity.
Python 3.9+, PyTorch, CLAP (Contrastive Language-Audio Pre-training), CLIP (Contrastive Language-Image Pre-training), AudioCLIP, PyWavelets (Discrete Wavelet Transform), scipy.signal (Butterworth Filters), torchgate (Spectral Gating), librosa (Audio Processing), torchaudio, Apache Parquet (Data Storage), Matplotlib, Seaborn, scikit-learn (t-SNE), NumPy, Pandas, HuggingFace Transformers, ONNX Runtime (Deployment), TensorRT (Optimization)