Research / Engineering / Experiments

Sound → Perception → Intelligence

Dahong Luo

01 / Selected builds

PROJECTS

01 / Experiment / 03.2025

Cat Diffusion
Model

A cat-image diffusion experiment built with PyTorch and a U-Net. Noise, structure, and the emergence of a recognizable form.

PyTorch / Diffusion / UNet

02 / Experiment / 03.2025

Quantum Neural
Network

A quantum–classical neural-network experiment combining PyTorch, Qiskit feature maps, variational circuits, and autoencoders.

PyTorch / Qiskit / ZZFeaturemap / RealAmplitude / Autoencoders

03 / Experiment / 03.2025

Fine-Tuning
SAM2

Fine-tuning the Segment Anything Model with PyTorch: an exploration of visual boundaries and image segmentation.

PyTorch / Segment Anything Model

04 / Experiment / 12.2024

Mathematical Figure
Transcriber VLM

A vision-language project for transcribing mathematical figures, using Llama and Unsloth on Google Cloud.

PyTorch / Unsloth / Llama / GCP

05 / Experiment / 10.2024

Embedded
Conditional GAN

An embedding-conditioned generative adversarial network, exploring how a learned condition can guide generation.

PyTorch / TCGAN / Embedding

06 / Experiment / 06.2024

Numpy Neural
Network

Neural-network and Transformer experiments in NumPy, with automatic differentiation as a building block.

NumPy / AutoGrad / Transformer

07 / Experiment / 05.2024

Cyberpunk 2077
Steam Review Analysis

An analysis of player reviews using language representations and machine-learning tools.

PyTorch / pandas / LightGBM / BERT

08 / Experiment / 01.2024

Reinforcement
Learning Pacman

A Pacman agent exploring reinforcement learning and Q-learning in a Pygame environment.

PyTorch / Reinforcement Learning / Q-Learning / PyGames

09 / IRONSIGHT Hackathon / 02.2026

AnchorVision

SLAM-based 3D object mapping from video

A 4D spatiotemporal mapping pipeline that combines visual SLAM and sparse UWB ranging to build persistent site memory without GPS.

ORB-SLAM3 / UWB / YOLO / LiDAR / C++ / ZeroMQ

  • Integrated ORB-SLAM3 with sparse Ultra-Wideband ranging to maintain a globally consistent 3D site memory.
  • Developed multi-agent alignment using UWB constraints as a global prior in a Maximum A Posteriori optimization solver, eliminating drift and enabling map fusion without visual loop closures or heavy overlap.
  • Implemented 2D-to-3D semantic lifting by projecting YOLO detections through ray-casting and LiDAR depth-map fusion, creating a persistent object index keyed by time and coordinates.
  • Built a C++ and ZeroMQ bridge to stream RGB-D and IMU data from iPhone LiDAR to a backend fusion service with sub-100 ms latency.

10 / HackUMBC / 10.2024

Grand Prize

Agent-Based
ELMS Copilot

A multi-agent educational assistant for syllabus parsing, course-content questions, and personalized quiz generation.

LangGraph / LangChain / Pinecone / PyTorch / Flask / React

  • Engineered specialized agents with LangGraph and LangChain for syllabus parsing, content QA, and automated quiz generation.
  • Architected a RAG pipeline using Pinecone vector databases for low-latency semantic search across heterogeneous course materials.
  • Developed a custom question-generation engine using PyTorch and OpenAI LLMs, fine-tuned to extract pedagogical concepts from notes and generate tailored assessments.
  • Improved retrieval using advanced chunking and metadata filtering to maintain context across complex subjects.
  • Built a Flask backend and React/Vite frontend with Tailwind CSS, integrating PyTorch for pedagogical content generation and student analytics.

11 / Hacklytics / 02.2024

1st Place — Elevance Healthcare Challenge

Medical Data
Compilation System

A multilingual medical-document platform connecting document ingestion with natural-language database retrieval.

LLMs / SQL / OCR / Translation / Document Parsing

  • Developed an upload pipeline that performs OCR, translation, and document parsing into SQL databases.
  • Built a retrieval pipeline that converts natural-language queries into SQL, summarizes results, and suggests diagnoses.
  • Created a web application for managing and analyzing medical documents in multiple languages and formats.

12 / HoyaHacks / 02.2024

1st Place — Best Digital Forensics Related Hack

Crime Event
Detector

An audio-driven video analysis system that finds critical acoustic events and turns long recordings into timestamped, searchable logs.

CNN / YAMNet-inspired Architecture / Audio Processing / MongoDB

  • Developed a CNN-based audio segmentation model inspired by YAMNet to detect events such as gunshots, car crashes, and screams in lengthy videos.
  • Curated custom datasets, applied data augmentation, and systematically tuned hyperparameters to distinguish target sounds from complex background noise.
  • Built a pipeline for audio extraction, temporal segmentation, and timestamped event logs to reduce manual review in forensic and law-enforcement applications.
  • Created a full-stack application with video upload, search, and MongoDB persistence for processed logs and detection results.

13 / Hack@CEWIT / 02.2024

1st Place — Spark of Genius Prize

Gesture Recognition
Control System

A multimodal interface connecting hand movement and voice to operating-system controls, a robotic claw, and an expressive air-guitar instrument.

OpenCV / Whisper / Hand Tracking / Robotics / Audio

  • Engineered real-time OpenCV detection and tracking to map hand kinematics into operating-system commands for hands-free navigation.
  • Integrated OpenAI Whisper for low-latency speech-to-text, allowing users to alternate between gestures and voice dictation.
  • Built a hardware bridge from virtual spatial coordinates to physical actuation, piloting a custom robotic claw through real-time hand tracking.
  • Developed continuous gesture-to-audio mapping based on hand positions, demonstrating responsiveness through an air-guitar interface.

02 / Research

RESEARCH

10/2025 — Present / Student Researcher

Neural Audio Codec

University of Maryland
Prof. Nirupam Roy · iCoSMoS Lab

A compact neural codec for real-time emergency voice transmission on low-power hardware.

  • Developing a compact neural audio codec optimized for low-power microcontrollers, enabling real-time audio compression and reconstruction.
  • Designing a low-power codec using residual finite scalar quantization and an EnCodec + Vocos-style model with < 1/20 the parameters of SOTA (~390K vs. 7M+), for emergency voice transmission in low-bandwidth, low-CPU settings: 2.3 kb/s and RTF 0.8 on a Raspberry Pi CPU.
  • Created automatic training scripts, implemented Weights & Biases logging, and performed ablation studies.

06/2025 — Present / Research Assistant

Perception-Aware Audio

University of Rochester
Prof. Zhiyao Duan’s Group

Psychoacoustic compression and query-guided source separation for higher-fidelity neural audio.

  • Integrated MP3/OGG psychoacoustic models to improve perceived quality in neural audio systems; analyzed LAME encoder stages and created Python bindings.
  • Bound LAME MP3 and OGG Vorbis psychoacoustic features to Python through the PyTorch C++ interface, combining them with a finite scalar quantizer and Transformer-based flow matching for perception-aware, high-fidelity neural audio compression.
  • Trained a U-Net-based joint audio separation model to distinguish multiple speakers and sources through natural-language queries. Led annotation and augmentation, iterated with PyTorch Lightning and Weights & Biases, and used grid search for hyperparameter tuning.

04/2025 — 09/2025 / Research Assistant

3D Human Reconstruction

University of Texas at Austin
Hezhen Hu, Postdoctoral Researcher · VITA Lab

Faster Gaussian-splatting reconstruction with measurable 3D mesh quality assessment.

  • Enhanced PGSR-based 3D human reconstruction with a Gaussian Splatting-inspired loss and a faster testing pipeline. Profiled training bottlenecks and reduced training time by 80%.
  • Introduced quantitative 3D mesh assessment using point clouds generated from output meshes and ground truth, automating validation and making output-quality tracking more reliable.

02/2025 — 09/2025 / Research Assistant

Research Writing Assistant

University of Maryland
Prof. Leo Zhicheng Liu’s Group

A retrieval-augmented workspace for literature review, scientific writing, and dataset exploration.

  • Reviewed literature on generative-AI research assistance, collecting and analyzing papers to identify trends and key methodologies.
  • Worked on a RAG-based writing assistant to help researchers write scientific papers.
  • Contributed to a project helping event-sequence data scientists process datasets.
  • Built metadata-driven graph re-ranking that improved retrieval accuracy, and integrated it into an interactive interface for faster evaluation.

09/2024 — 12/2024 / Student Researcher

Sound Event Detection

University of Maryland
LSTM Sound Event Detection Research · Honors Seminar

An efficient causal LSTM detector for real-time acoustic event recognition from waveforms.

  • Conducted an Honors Seminar research project and wrote a paper as first author.
  • Developed an LSTM-based sound event detection model that detects events from waveforms without relying on future data.
  • Conducted an ablation study to better understand the model and improve it efficiently.
  • Achieved real-time inference at 462 ms per 10-second clip with ~80% accuracy, compared with non-real-time CNN-LSTM state-of-the-art models.
  • Later extended the system with a pretrained RAVE encoder and the Mamba temporal model, improving accuracy by 10% while preserving efficiency.

03 / Work experience

EXPERIENCE

Boston Consulting Group

San Francisco, CA

  • Developed Unified Benchmarking Database, consolidating heterogeneous benchmarking datasets from across BCG into a standardized database and deploying an agentic retrieval layer for efficient access to benchmarking data.
  • Onboarded 4 new benchmarking databases by building reusable ingestion scripts and an automated data-ingestion quality testing pipeline, while contributing to the agentic retrieval system and end-to-end data integration workflow.

University of Maryland, Engineering IT

College Park, MD

  • Developed a Flask/React platform for university makerspaces to manage equipment access, enforce safety permissions, and track machine usage through secure authentication and streamlined workflows.
  • Built a Gemini-based agentic code review system that retrieves relevant code via snippets and CLI tools, evaluates implementations against department policies, and aggregates outputs from three task-specific models with a fourth validation model, reducing review time by 30%+.

IZAI Corp

Tokyo, Japan

  • Built and deployed a Flask/React RAG application on Azure using BM42 Hybrid Search and LlamaIndex; developed a metadata-driven agentic graph retrieval system with multi-reranking for efficient data retrieval.
  • Fine-tuned MetaVoice TTS on ReazonSpeech with dataset cleaning and augmentation, and implemented 8-bit quantization optimizations that improved generation speed by 2×+ while maintaining audio quality

InfoDeliver Corp

Tokyo, Japan

  • Developed an LLM-based agentic automation system using GPT-4 Vision to perceive UI elements and autonomously execute computer actions; built a real-time Whisper voice interface for natural-language task control and mid-task adjustments.
  • Created and annotated a web UI component dataset to fine-tune YOLOv7, achieving 70% detection accuracy, and integrated the detector into a modular perception, chain-of-thought, and computer-action pipeline.

04 / Education & toolkit

EDUCATION

2023 → 2026 / Tokyo → Maryland → New York

  1. —

    University of Tokyo

    Tokyo, Japan

    B.S. studies in Natural Sciences · Subsequently transferred to the University of Maryland

  2. —

    University of Maryland

    College Park, Maryland

    B.S. · Computer Science major · Mathematics minor

    GPA 3.90 / 4.0

  3. —

    Columbia University

    New York

    M.S. Electrical Engineering

Interests & skills

Interests

Audio Processing · Signal Processing · Generative AI · Large Language Models · Computer Vision · Natural Language Processing · Quantum Machine Learning

Languages

Python (7+ years) · C · C++ · Java · JavaScript · MIPS Assembly

Machine learning

PyTorch · TensorFlow · Keras · Qiskit · PennyLane · NumPy · scikit-learn · JAX · pandas · MATLAB · seaborn · Transformers · YOLO · Reinforcement Learning · Llama · LangChain · Unsloth

Technology stack

Azure · Docker · Git · Flask · REST APIs · React.js · Next.js · MongoDB · SQL · Hugging Face