Career trace — reverse chronological
Freelance — Secured Globe (US-based client)
Engineered a specialized knowledge extraction pipeline to isolate a target 8% subgraph from the complex Foundational Model of Anatomy (FMA) medical ontology. The primary technical challenge was maintaining full semantic integrity and logical reasoning context during graph slicing. To achieve this, authored complex SPARQL queries capable of correctly parsing blank nodes and navigating intricate OWL existential restrictions (owl:someValuesFrom). To address performance bottlenecks, migrated the graph traversal logic from native Python to the high-speed Rust-powered Oxigraph engine, followed by rigorous validation to ensure zero structural or semantic degradation in the extracted subset.
Specializing in adapting and fine-tuning state-of-the-art (SOTA) ML/DL architectures for complex Computer Vision, Signal Processing, and 3D modeling challenges. Pairing deep algorithmic expertise with rigorous validation protocols to extract maximum performance from models. Consistently proving solution reliability and top-tier global execution on Kaggle, achieving Top-2% and Top-5% ranks globally.
Competed in a digital archaeology challenge organized by a Chinese archaeological institute, tasked with virtually reconstructing original vases from a dataset of over 17,000 ancient pottery shards excavated from a Shang Dynasty burial site. Following an extensive review of existing literature and algorithmic solutions, which proved inapplicable to the dataset's unique constraints, engineered a novel reconstruction pipeline from scratch. The deployed solution utilized an advanced clustering approach that integrated classical computer vision features, texture analysis, and Self-Supervised Learning (SSL) embeddings to accurately group related fragments. Delivered a comprehensive technical presentation detailing the feature extraction methodology and clustering approach for the archaeological community.
Finding pre-Columbian villages on satellite images, with validation in textual sources
Engineered a visual assessment system utilizing multi-channel satellite imagery to estimate the probability of pre-Columbian settlements, successfully discovering the two most relevant uncharted sites during an OpenAI-sponsored Kaggle competition. The overall objective required leveraging diverse open-source data — including LiDAR, multispectral satellite imagery, and historical texts — alongside LLM models to identify hidden archaeological traces beneath the Brazilian Amazon canopy. Acting as the sole technical specialist on a cross-functional team, established the technical methodology and analytical pipeline, strategically targeting a high-risk area designated for dam construction to uncover endangered archaeological sites before their potential destruction. The implemented solution involved biological zoning to isolate flora historically utilized by pre-Columbian populations, effectively filtering out post-expansion species, plus multispectral satellite image analysis to detect potential geoglyphs, and collaboration with a domain expert to optimize a prompt for assessing settlement probability. The project culminated in a comprehensive public report detailing the discovered settlements and methodological approach.
kaggle.com/competitions/openai-to-z-challenge/writeups/lost-city-of-z
Secured a Silver Medal (29th place globally) in the 2025 Stanford RNA 3D Folding Kaggle competition, focusing on the prediction of complex three-dimensional RNA structures. To tackle this highly specialized challenge, strategically partnered with a bioinformatics expert, forming a cross-functional team that perfectly balanced deep biological domain knowledge with advanced machine learning capabilities. As the primary ML engineer, was responsible for adapting, deploying, and rigorously evaluating state-of-the-art neural network architectures from leading research laboratories within the constrained Kaggle environment. Technical contributions included extensive experimentation with cutting-edge models, notably engineering a LoRA fine-tuning pipeline for an open-source implementation of AlphaFold3. Through rigorous validation, determined that architectures incorporating cross-species biological data yielded significantly superior predictive performance — a data-driven insight that allowed the team to strategically pivot away from the AlphaFold3 approach, optimizing the final ensemble to achieve a top-tier global ranking.
Secured a Bronze Medal (103rd place globally) in the 2024 RSNA Lumbar Spine 3D Degenerative Classification Kaggle competition by developing a computer vision solution to detect lumbar nerve impingement from MRI scans. Collaborating closely with a teammate, took charge of validating and debugging a complex multi-stage model architecture designed to process sequences of 2D slices into cohesive 3D anatomical representations. The pipeline consisted of segmentation of target vertebrae on vertical slices, algorithmic cross-referencing to match horizontal slices to detected landmarks, and precise classification of impingement zones. Once the baseline framework was fully validated, the team strategically divided research efforts to rapidly iterate and test independent hypotheses in parallel. Core technical contributions included significantly improving the public algorithm for spatial mapping between horizontal and vertical slices, alongside heavily optimizing a YOLO-based detection baseline. By successfully merging parallel modeling efforts and insights, the team engineered a highly accurate ensemble solution that resulted in a top-tier global finish.
Developed a system to determine the position and shape of a medical device on intraoperative medical images. The primary objective was to automate device detection, identify its edges, and estimate its shape and spatial orientation — tasks previously performed visually by the surgeon. As the problem resided at the intersection of 2D and 3D computer vision, initial experiments with classical computer vision techniques proved ineffective due to high visual variability, even on controlled laboratory images. To overcome this, a multi-stage machine learning pipeline was engineered: YOLOv5 for initial object detection, followed by precise segmentation of the device and its edges, plus a custom loss function to calculate the most probable spatial alignment of the object. The resulting system surpassed human-level efficiency and accuracy. It is currently deployed as a surgical assistant tool, pending extensive regulatory approval before it can be authorized to make autonomous surgical decisions.
Medical startup
Imperial College London
Wikidata-like Wikibase instance
Transformed a locked geographic and cultural ontology into a web-accessible Wikibase Knowledge Graph to empower collaborative research. Established internet routing for the initial local data storage, bootstrapped a Dockerized Wikibase stack, and engineered a Python-based ETL pipeline. The custom script processed structured CSV exports and converted them into semantic triplets, populating the Wikibase repository with fully queryable entities and properties.
Open Data Science — taught by Mikhail Galkin, senior researcher at Google
For Telegram channel administration and customs sales
Summer school in Computational Neuroscience
Faculty of Psychology, Department of Psychophysiology · Area of studies: Computational Neuroscience
Engineered a machine learning pipeline to classify electroencephalographic (EEG) recordings for the detection of intentionally concealed information, leveraging the p300 cognitive evoked potential. Adapted and applied modern Brain-Computer Interface (BCI) algorithms, specifically those utilizing Riemannian geometry developed by A. Barachant, to a novel psychophysiological domain. Developed the complete data processing and evaluation workflow in Python utilizing mne, pyriemann, and scikit-learn. The technical pipeline involved frequency filtering (1–20 Hz), epoch extraction, and mitigating inherently low signal-to-noise ratios through targeted epoch averaging. For feature extraction and classification, covariance matrices were computed from the EEG channels, projected onto a Riemannian manifold to reduce multi-modal noise, and rigorously evaluated Minimum Distance to Mean (MDM) and Logistic Regression classifiers against a Common Spatial Patterns (CSP) baseline.
Worked on an fMRI scanner with a team of psychologists, participated in composing psychological test batteries, and in running and processing MRI studies.
Incomplete higher education — Control Systems for Aircraft