Dan Ben Ami

About

AI researcher specializing in deep learning, computer vision, and vision-language models (VLMs).

I design and build innovative, scalable AI systems and complex pipelines, integrating creative problem-solving with expertise in multimodal learning, algorithm optimization, and end-to-end system design for real-world impact.

Work Experience

Computer Vision Data Scientist, PhD Intern

Microsoft · WSSI · Nov 2025 – Present

  • Engineered edge-constrained image restoration pipelines, adapting SOTA backbones (NAFNet, HAT) via Knowledge Distillation to bridge the gap between high-capacity teacher models (+3.5 dB PSNR) and strictly bounded edge architectures.
  • Formulated domain-adaptive distillation criteria and hybrid loss objectives, dynamically coupling teacher supervision with ground-truth signals to achieve a +0.7 dB PSNR boost under fixed latency and memory footprints.
  • Designed auxiliary loss functions targeting localized reconstruction artifacts (chessboard, ringing) during denoising and super-resolution, significantly boosting perceptual quality and Mean Opinion Scores (MOS).

Example of denoising with the off-the-shelf NAFNet model (Chen et al., ECCV 2022), as my own work's data is confidential.

Denoising example using NAFNet

Computer Vision & AI Algorithm Developer

Elop · Feb 2023 – Oct 2025

  • Performed a diverse range of Computer Vision and Vision-Language tasks and developed a deep understanding of various state-of-the-art (SOTA) models and architectures.
  • Received project requirements and system constraints, then designed, optimized, and integrated custom algorithmic pipelines into larger systems.
  • Conducted extensive literature reviews—evaluating over 70 research papers annually on customized data with numerical performance metrics.

Scene Understanding and Threat Recognition System

  • Built an end-to-end pipeline that ingests multi-object tracking outputs and leverages Video-LLMs for complex-scene understanding and per-object threat scoring.
  • Led a 3-person project; authored and executed a one-year roadmap (versioning, scope per release, task allocation) to deliver the project.
  • Benchmarked multimodal models from 3B–78B (Qwen2.5-VL, LLaVA-OV, etc.) on multiple datasets, comparing accuracy and runtime/latency to select deployment candidates.

Multi-Object (Multi-Sensor) Tracking

  • Developed a tracking solution that increased precision from 72% (naive off-the-shelf YOLO+DeepSORT pipeline) to 93% using a customized pipeline.
  • Integrated an ensemble of video classification models (MViT, TimeSFormer, etc.) focused on reducing false alarm rates.
  • Implemented a three-sensor tracking system (RGB, SWIR, MWIR) to ensure robust performance across different imaging modalities.

Visualization of simple basic off-the-shelf multi-object tracking in RGB (Vis) sensor without the full pipeline (due to IP and confidentiality).

Zero-Shot Keypoint Tracking

  • Explored zero-shot tracking for keypoints in infrared videos characterized by extreme noise conditions.
  • Utilized simulative data generated with UE5 specifically for testing under high interference scenarios.
  • Rigorously evaluated SOTA models to determine limitations and ensure their suitability under adverse conditions.

Visualization of tracking keypoints on a passenger airplane from standard RGB (no noise) video (due to IP and confidentiality).

Video Restoration and Super Resolution

  • Investigated and compared multiple SOTA models, including a self-designed architecture and a self-trained model.
  • Applied knowledge distillation from a FeMaSR teacher model to optimize the chosen architecture.
  • Achieved performance improvements of 5–20% over classical methods, as measured by metrics such as MTF, SSIM, and ESF.

Demonstration of video restoration and turbulence mitigation. This output is from the DATUM model (Zhang et al., CVPR 2024), as my own work's data is confidential.

Video Classification for False Target Filtering

  • Designed a comprehensive data management infrastructure for experimental model training and evaluation.
  • Developed tools for dynamic data integration, version management, attribute-based filtering, and complex augmentations (e.g., smart copy-paste and 3D rotations relative to the camera plane).
  • Conducted hundreds of training experiments across both lightweight and heavy SOTA architectures, rigorously evaluating each component (data version, augmentations, architecture, transfer learning, etc.).
  • Increased overall accuracy by 22%.

Teaching Assistant

Ben-Gurion University · 2020 – 2023

  • Algorithms and Graph Theory (2023)
  • Digital Design (2022)
  • Digital Computers Structure (2022)
  • Linear Algebra (2020–2021)

Publications

HiMu pipeline

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering

Dan Ben-Ami, Gabriele Serussi, Kobi Cohen, Chaim Baskin

Under review2026

We introduce HiMu, a training-free hierarchical multimodal frame selection framework for long-video QA that decomposes queries into logic trees with specialized experts, advancing the efficiency–accuracy Pareto front while requiring ~10× fewer FLOPs than agentic systems.

HERBench teaser

HERBench: A Benchmark for Multi-Evidence Integration in Video Question Answering

Dan Ben-Ami, Gabriele Serussi, Kobi Cohen, Chaim Baskin

CVPR2026

We present HERBench, a VideoQA benchmark for multi-evidence integration across time, requiring aggregating multiple non-overlapping evidential cues. Evaluating 13 state-of-the-art Video-LLMs reveals pervasive failures (31–42% accuracy vs. 20% random baseline), disentangled into retrieval and fusion deficits.

Spectrum
Access

A Stable Polygamy Approach to Spectrum Access with Channel Reuse

Dan Ben Ami, Kobi Cohen

Preprint2024

I introduced the "Stable Polygamy Problem" (SPP) for spectrum access with channel reuse, developed efficient algorithms including RP&R, and proved their performance in specific interference regimes with strong simulation results.

Video-Text
Retrieval

VEDJE: Video-Efficient Discriminative Joint Encoder for Scalable Video-Text Retrieval

S. Wagner, G. Serussi, D. Ben-Ami, T. Galanti, C. Baskin

ICML WorkshopAdaptFM: Resource-Adaptive Foundation Model Inference

An efficient joint video-text encoder that makes discriminative video-text retrieval scalable under tight inference budgets.

Federated
Learning

PAUSE: Low-Latency and Privacy-Aware Active User Selection for Federated Learning

Ori Peleg, Natalie Lang, Dan Ben Ami, Stefano Rini, Nir Shlezinger, Kobi Cohen

IEEE Trans. Signal ProcessingVol. 73, 4556–4572, 2025

PAUSE jointly addresses privacy-leakage accumulation and communication latency in federated learning via a multi-armed bandit algorithm for active user selection, with a theoretical reward-growth analysis and a simulated-annealing relaxation for reduced complexity.

Gene
Expression

A Universal System for Boosting Gene Expression in Eukaryotic Cell-Lines

Inbal Vaknin, Or Willinger, Jonathan Mandl, Hadar Heuberger, Dan Ben-Ami, Yi Zeng, Sarah Goldberg, Yaron Orenstein, Roee Amit

Nature Communications2024

Designed and trained deep learning models to predict protein expression from DNA sequences, guiding motif selection for cross-species promoter boosting.

Education

Ph.D. in Electrical & Computer Engineering

Ben-Gurion University · 2024 – Present

Lab: InsightLab
Advisors: Dr. Chaim Baskin, Prof. Kobi Cohen

Combined Track for top-performing students.

M.Sc. in Electrical & Computer Engineering

Ben-Gurion University · 2021 – 2023

Advisor: Prof. Kobi Cohen
GPA: 97
Honors: Summa Cum Laude

Accelerated Direct M.Sc. program for top-performing students.

Main courses: Deep learning, Sequential learning, Statistical inference and Data Mining, Multivariate statistical data analysis, Game theory.

B.Sc. in Computer Engineering

Ben-Gurion University · 2018 – 2022

GPA: 92
Honors: Summa Cum Laude (dean's list)
Ranking: Ranked #1 in my class (2018–2022 BGU CE)

Research: Computational modeling of protein-DNA/RNA interactions using deep learning models, guided by Dr. Yaron Orenstein.

Programming

Programming Languages
Python, C++, MATLAB, Bash
Deep Learning Frameworks
PyTorch, TensorFlow, Keras
Computer Vision & VLMs
OpenCV, Hugging Face Transformers, vLLM, CLIP, LLaVA, Qwen, SAM, YOLO
Data Processing & Analysis
NumPy, Pandas, scikit-learn, SciPy
Visualization
Matplotlib, Plotly, TensorBoard
Optimization & MLOps
PyTorch Lightning, Hydra, ONNX, Torch-TensorRT
Tools & Environments
Docker, Git, Linux, Conda, Jupyter

Soft Skills

Creative problem-solving and innovation in AI system design Strong analytical and critical thinking abilities Cross-disciplinary collaboration and communication Rapid learning and adaptation to emerging technologies Translating research concepts into practical, scalable solutions Project planning and execution in complex, multi-stage pipelines Mentoring and knowledge sharing in technical teams Attention to detail while maintaining big-picture perspective

Contact

Email: danbenami3@gmail.com

LinkedIn: Dan Ben Ami