I design and build generative-AI services, conversational assistants, and end-to-end data pipelines on Azure and GCP — contract-first, with deterministic validation bounding what the model is allowed to decide, and formal evaluation against traceable ground truth. Published researcher in hybrid RAG, graph-based fraud detection, and applied causal inference. Nearly three years of formal line management before specialising.
A causal-inference evaluation of Peru's 2024-2026 states of emergency, combining a staggered difference-in-differences design (Callaway-Sant'Anna) with partial-identification bounds (Manski) over an administrative district-month panel (SIDPOL, 1,839 districts), emergency decrees, and the ENAPRES victimization survey. Shows that much of the official crime trend reflects endogenous underreporting rather than a real change in crime, and that the sign of the effect is not point-identified under credible assumptions.
An exploratory pilot comparing three LLM agent deployment patterns — Managed Agents, a DIY multi-agent factory, and Direct API — on financial data governance tasks. Finds that platform-level prompt caching can invert the expected cost hierarchy: an uncached DIY loop accumulates up to 21.68x more context tokens than a single-pass baseline, while managed caching cuts repeated-context cost by roughly 90%. Includes a security/data-sovereignty analysis for regulated fintech (PCI-DSS, SBS Peru, Ley 29733).
Introduces the GRAFID framework with three novel metrics (Feature Richness Index, Graph Signal Gain, Cost-Effectiveness Ratio) to determine when GNNs outperform XGBoost for fraud detection. Evaluated on IEEE-CIS (590K transactions) and Credit Card EU (284K) datasets with 20+ model configurations, non-parametric statistical validation (Wilcoxon, McNemar), and multi-seed reproducibility.
Compares four retrieval strategies — hybrid search with RRF, HyDE, multi-query expansion, and lightweight graph-based retrieval — on a single PostgreSQL + pgvector platform, with no dedicated vector database. On a 427-document Spanish corpus (22 ground-truth pairs, 32 evaluation runs), HyDE reached the strongest scores (P@5: 0.595, MRR: 0.838, Hit Rate: 1.0), a preliminary ~23% relative gain over the hybrid baseline pending larger-sample validation.
Design and prototype of an autonomous two-wheeled mobile robot using inverted pendulum control systems, with stress simulations and material validation in Autodesk Inventor.
Applied AI engineering: turning business ideas into production services, on Azure and conversational platforms.
Distributed architecture for natural language query (NLQ) processing translating business questions into SQL. Asynchronous WebSocket orchestration with adaptive timeouts (5s→40s), semantic engine using open-source LLMs for NLQ-SQL with ontological schema mappings (50+ business terms), hierarchical fallbacks, and conversational interface on WhatsApp. Democratized BI access for non-technical enterprise users.
A from-scratch forecasting system for the 2026 World Cup knockout stage, trained and run entirely on local hardware with no cloud infrastructure. Combines a Dixon-Coles bivariate Poisson scoreline model (attack/defense strength, home advantage, time-decayed match weighting, low-score correlation correction) with pre-match Elo ratings and a Monte Carlo bracket simulator (extra time & penalties) for tournament-winner probabilities, optionally blended against bookmaker market odds. Backtested out-of-sample on all 72 group-stage matches the model never trained on: 62.5% match-winner accuracy (on par with a pure Elo baseline) and 12.5% exact-scoreline accuracy. A calibration analysis found the model measurably underconfident (optimal temperature ≈ 0.80) — but with only 72 held-out matches a global recalibration would overfit, so that was documented as a diagnostic finding rather than deployed live.
Open to collaborations, research partnerships, and new opportunities.