Building systems that reason at scale
From writing business rules on enterprise stacks to orchestrating autonomous AI agents under regulatory constraints — each phase built the intuition the next one needed.
Five phases. One compounding arc.
Data Scientist and applied ML researcher with 15 years building production AI systems at the intersection of large language models, graph analytics, and enterprise compliance. Currently architecting multi-agent LLM orchestration and knowledge graph infrastructure for regulatory workflows in financial services — and independently researching in-context learning theory and model compression. I write about the gap between ML research and production reality: what the papers don't tell you, what breaks at scale, and what's worth building from scratch. I'm also the creator and maintainer of northgradient.com — a free, interactive platform teaching how AI actually works from first principles.
Below: the arc that got me here — not a career history, but a sequence of transformations where each identity was prerequisite to the next.
Selected Projects
NorthGradient: Free, Interactive AI Education
A free, ad-free library teaching how ML & LLMs actually work from first principles — full courses, short reads, and guided paper walkthroughs. Designed, built, and kept running end to end by me: static, edge-served on Cloudflare, with auth, highlights, notes, and spaced-repetition review.
Cardinal: Synthetic Credit-Card Data
A conversational interface turns requirements into validated specifications, dependency graphs, and reproducible credit-card data samples. Combines an ontology-guided builder with a deterministic simulation engine using demo parameters.
MetaModel: Code-Derived Data Lineage
Builds a Neo4j knowledge graph from Python code to explore data dependencies, transformations, and business rules. Distinguishes statically extracted facts from LLM-inferred information.
ReconX: Regulatory Reconciliation
Compares source and target reporting datasets, calculates discrepancies, and uses domain-guided LLM classification to explain potential causes and remediation. Includes a React interface for inspecting workflow progress.
efors: Research Paper Discovery
Collects research papers from arXiv, Semantic Scholar, and OpenAlex, deduplicates records across sources, and ranks papers using configurable scoring signals. Provides a FastAPI backend and interfaces for exploring the collection.
Multi-Agent Orchestration at Scale
Distributed agent system on AWS coordinating specialized LLMs across Neo4j knowledge graphs and Bedrock inference for enterprise compliance workflows.
AI-Curated Data Quality Governance at Scale
AI-enabled system that ingested, curated, and refined data quality rules across disparate source systems — translating heterogeneous rule definitions into standardized, configurable Great Expectations YAML for enterprise-wide validation.
Structural Entropy in Fraud Networks
Neo4j-based fraud detection using community detection (Louvain, k-core, label propagation) and structural entropy measures across transaction networks.
Agentic Regulatory Control Extraction
Decomposed federal deposit insurance mandates into 110 structured IT controls using an 8-pass LLM extraction pipeline — each control anchored to its source regulatory provision, enabling automated gap analysis and audit-ready compliance traceability.
Log Mining to Enterprise Event Platform
Mined and curated 12+ security event types from raw, unutilized enterprise logs — evolving into a no-code/low-code detection platform that centralized event observability org-wide via YAML-driven pattern configuration.
Iterative Relation Extraction with LLMs
Multi-pass relationship extraction system using the Claude API with pronoun resolution and iterative refinement across unstructured text corpora.
Research Interests
In-Context Learning Theory
Formalizing ICL success conditions through an information-theoretic lens — mutual information between prompt features and output correctness, tractable bounds, and empirical predictions.
Memorization in Transformer Layers
Empirical study of how memorization concentrates across transformer layers — identifying which layers disproportionately contribute to verbatim recall using GPT-2 on A100 GPU.
Model Quantization & Compression
Hands-on exploration of post-training quantization on 7B-parameter models — per-layer error accumulation, accuracy/compression tradeoffs, and what GPTQ actually looks like from scratch.
Tech Stack
- Claude (Anthropic)
- AWS Bedrock
- LangChain
- MCP Protocol
- Hugging Face
- GPTQ / bitsandbytes
- Neo4j + GDS
- Apache Spark
- Kafka
- Great Expectations
- Oracle SQL
- Cytoscape.js
- AWS EC2 / EKS
- Docker / Kubernetes
- Python / Scala
- Streamlit
- FastAPI
- Git / GitHub
- PyTorch
- A100 GPU (CUDA)
- arXiv
- Semantic Scholar API
- GPT-2 / 7B Models
- LaTeX
Let's connect
Open to research collaborations, technical discussions, and consulting on LLM systems, graph analytics, or compliance AI in financial services.