Research Projects
GazeEval-VLM 2026
Independent project building a lightweight framework to test how human gaze, used as an external cognitive signal, affects vision-language models on visual and chart question answering. Model under test: Qwen2.5-VL-3B, benchmarked on the VQA-MHUG and ChartGaze datasets.
Building GazeEval-VLM, a model-agnostic evaluation pipeline, and GazeVLM-Lite, a lightweight inference method that feeds gaze in as visual evidence through gaze cropping, gaze highlighting, background blurring, and multi-view gaze-guided inputs, without touching model weights. Running a cross-domain comparison across VQA and CQA subtasks to see which ones depend on local visual evidence versus global context, and where gaze guidance actually helps versus where it just adds noise. Also defining a gaze efficiency score to check whether the accuracy gained from adding gaze signal is worth the extra attention cost it introduces, rather than just assuming more signal is better.
Running the whole pipeline solo: gaze data preprocessing, the VLM inference pipeline, implementing each gaze-guided intervention method and its subdataset, the cross-domain comparative study, and the technical writeup.
Tech stack: PyTorch, Hugging Face Transformers, Accelerate, NumPy/Pandas, scikit-learn, Matplotlib, Jupyter Notebook, Gradio for result demos.
Dist-AI 2025
Independent open-source project simulating a distributed AI inference architecture with Python and Docker: a coordinator service dispatching requests to three worker containers, each wrapping a different pretrained model, to explore load-based routing and fault tolerance in a multi-model serving setup.
Routes each request by modality: text-only to a BERT worker (prajjwal1/bert-tiny), image-only to a MobileNetV3-Small classifier, and text+image to a CLIP worker (openai/clip-vit-base-patch32). Each worker runs in its own container behind docker-compose, with a per-worker health-check endpoint, retry logic, and structured error logging so one worker crashing doesn't take the others down with it.
Wrote a smoke-test script for single-request checks and a batch-test script that fires 21 concurrent requests to stress the routing logic, plus simulated worker failures to verify the retry and logging paths actually fire rather than assuming they work from reading the code.
Tech stack: Python, Docker, Docker Compose, Flask, REST APIs, PyTorch, Hugging Face Transformers, torchvision, Requests.
ValveSense-V2 2024
Independent project: a CEEMDAN-SVR forecasting system for industrial valve pressure sensor data, built to handle non-stationary signals that standard time-series models struggle with.
Pipeline decomposes the raw signal with CEEMDAN into IMF (intrinsic mode function) components, forecasts each with SVR using an RBF kernel over sliding windows, then reconstructs them into a final forecast. Split the code into separate decomposition, feature engineering, forecasting, and evaluation layers so each stage can be swapped or benchmarked on its own.
Wrapped the pipeline in a Gradio app, containerized it with Docker, and built a benchmark harness to compare forecasting configurations. Handled the full pipeline solo, from data preprocessing to deployment. LSTM and Transformer forecasting are next on the roadmap.
Tech stack: Python, NumPy, Pandas, Matplotlib, scikit-learn, PyEMD, Gradio, Docker.
Analysing Human vs. Neural Attention in VQA 2023 – 2024
Master's thesis at the Universität Stuttgart, supervised by Yao Wang and Susanne Hindennach.
Goal: extract attention maps from a Visual Question Answering model, compare them against real human eye-gaze data, and see whether the two actually line up or whether "neural attention" is more of a metaphor than a mechanism.
Re-implemented MCAN (Modular Co-Attention Networks) from scratch and matched published accuracy on VQAv2 within a point (67.16 vs. 67.17 overall for the small variant), using both Faster R-CNN region features and ResNet-50 grid features as separate image encodings. Extended the same setup to the GQA dataset to run a second, independent comparison, again landing close to the published OpenVQA baseline.
Built the pipeline to pull raw 1D attention weights out of the model, remap them into 2D heatmaps aligned to the original image, and apply Gaussian smoothing over region-based attention so it's directly comparable to the smoother human gaze heatmaps. Ran this against two human-gaze-annotated benchmarks, VQA-MHUG (3,990 question-image pairs, gaze averaged over 3 annotators per question) and AiR-D (1,454 pairs built on GQA), and scored the machine-vs-human overlap with AUC, Spearman's rank correlation, and Jensen-Shannon divergence.
Result: machine and human attention correlated significantly on both datasets (Spearman's ρ up to 0.68), with Gaussianized region attention consistently outperforming raw region or grid attention on AUC, suggesting the correlation is genuine and not an artifact of how the attention map is post-processed.
Tech stack: Python, PyTorch, torchvision, NumPy, SciPy, scikit-learn, Git for experiment tracking, Jupyter Notebook for analysis and visualization.
LatentFit 2023
Independent project exploring metric learning on Fashion-MNIST: a CNN embedding encoder trained with a triplet network to map fashion images into a 128-dimensional semantic space, rather than framing it as classification.
Built the triplet dataset construction, trained with triplet margin loss, and evaluated the resulting embedding space with t-SNE and cosine-similarity image retrieval. The t-SNE projection showed semantically similar items clustering together (sneakers near sandals, shirts overlapping pullovers), which the retrieval demo confirmed by pulling visually and semantically related items for a given query image.
Tech stack: Python, PyTorch, torchvision, NumPy, scikit-learn, matplotlib.
Visualization for Computational EEG Methods 2022
Seminar project during the master's program: built interactive Julia Pluto notebooks to visualize and explain core EEG preprocessing methods, ICA, baseline correction, and re-referencing, so the effect of each step on the signal is visible rather than just described.
Tech stack: Julia, Pluto Notebook, DSP.jl, Plots.jl.
Computational Modeling of Visual Saliency 2017 – 2019
Laboratory research project on computer vision and image saliency detection, spanning sophomore to senior year and culminating in my bachelor thesis under Prof. Dr. Bing Zeng. Built up computer vision fundamentals, then investigated how traditional and learning-based saliency models identify and prioritize informative regions in natural images, and where each family of approaches breaks down on interpretability or compute cost.
Proposed a saliency detection method built directly on image bitmaps and implemented it as a lightweight ML model, then benchmarked it against existing traditional and learning-based saliency models on runtime, memory footprint, and mean absolute error. The bitmap-based approach held up competitively against heavier models while using substantially less compute, which mattered for deploying saliency detection in resource-constrained settings.
Tech stack: MATLAB, MATLAB Image Processing Toolbox, Python, OpenCV, NumPy, scikit-image, SVM/Random Forest classifiers, classical CV techniques (edge detection, superpixel segmentation, frequency-domain analysis).