← Back to this milepost on the trail
Milepost 4 of 4 on the climb
Visa
- Role
- Data Science Intern
- Season
- May 2026 – Present
- Location
- Foster City, CA
- Field kit
- PyTorch · Llama · LoRA · Spark · Kubernetes · Hive
Recommending merchants to cardholders is a language problem disguised as a payments problem. I extended Llama with 10,000+ new merchant tokens, domain-specific embedding initialization, and LoRA fine-tuning to build a merchant recommendation model reaching 84% Recall@10 and 88% NDCG@10. Feeding it meant engineering a distributed Spark ETL pipeline on Kubernetes that turns 76+ billion raw transactions into natural-language training data, then 210+ ablation studies to find what actually mattered, ending 6x above the strongest heuristic baseline on unseen merchants.