Jesús Eduardo Rodríguez Saucedo

AI Engineer · Agentic Systems & Multi-Agent Orchestration · LLM Evaluation

01

Profile

AI Engineer with 10+ years building production-grade software and real-time AI systems, focused on agentic AI, multi-agent orchestration, and LLM-powered applications. I design systems where multiple specialized agents collaborate under a supervisory architecture — with structured outputs, tool integration, and auditable execution — and bring the same rigor to evaluating them: controlled baselines, repeated trials, ablations, and blind labeling to separate genuine model behavior from execution and prompt-scaffolding artifacts. Comfortable owning a system end-to-end, from asynchronous backend architecture to the evaluation harness that proves it works.

02

Core expertise

Agentic AI & Multi-Agent Systems
Multi-agent orchestrationsupervisor architecturesspecialized agentstool/function callingstructured outputsagent workflowsmulti-provider LLM systems
AI Engineering & Architecture
PythonFastAPIasyncioREST APIsWebSocketsdistributed & event-driven systemsreal-time AIAPI integrations
LLM Evaluation & Safety
Prompt engineeringreproducible evaluation designred teamingbehavioral robustnesshallucination detectionrisk classification
LLMOps & Observability
LangfuseDatadogstructured loggingexecution tracingevaluation pipelinesexperiment provenance
Cloud & Infrastructure
Google Cloud (Cloud Run, Cloud Build)Firebase/FirestoreDockerPostgreSQLRedisSupabase
Delivery
Git/GitHubGitHub ActionsCI/CDTypeScriptReactNext.js
03

Professional experience

Lead AI Engineer — VocalisAI

2025 – Present
Independent / Google Cloud
Finalist — Google Gemini Live Agent Challenge 2026
  • Architected a production-oriented multi-agent voice intelligence platform: one supervisor agent orchestrating six specialized agents (reception, qualification, emergency triage, billing, follow-up, outbound) on Gemini Live, FastAPI, WebSockets, Twilio, and Google Cloud.
  • Built real-time evaluation for hallucinations, boundary violations, and false-urgency behavior, with corrective intervention and hard-veto mechanisms for high-risk model actions.
  • Designed low-latency asynchronous execution pipelines for real-time voice interactions, with structured validation and observability built into the agentic pipeline.

AI Systems Architect — Multi-LLM Evaluation

2024 – Present
Independent
Finalist — Google Cloud AI + Datadog Hackathon
  • Designed multi-provider LLM orchestration and evaluation architectures coordinating several providers through structured, comparative workflows.
  • Implemented structured outputs, stage-level observability, execution checkpoints, and auditable experiment artifacts.
  • Built evaluation pipelines supporting repeated trials, baseline comparison, component ablations, and behavioral analysis, with infrastructure to separate model behavior from execution and parser artifacts.

AI Systems Engineer — LeadHunter

2025 – Present
Independent
  • Built an AI-powered lead qualification pipeline: Normalize → Deduplicate → LLM Scoring → Drafting → CRM, integrating PostgreSQL/Supabase, HubSpot APIs, and Langfuse observability.
  • Implemented structured validation and deterministic processing stages around probabilistic LLM components, turning unstructured lead data into validated CRM-ready records.

Lead Developer & Solutions Architect

2020 – 2024
World Vision Telecom
  • Engineered production-grade real-time communication systems supporting 500K+ interactions, using Python, FastAPI, Redis, and WebSockets on cloud infrastructure.
  • Built asynchronous processing pipelines for high-volume communication workloads, with emphasis on reliability, scalability, and operational resilience.
  • Led architecture and implementation decisions across backend services and real-time application infrastructure, integrating external APIs and communication platforms into production workflows.

Software Engineer

2014 – 2020
Earlier Experience
  • Software engineering, systems architecture, web platforms, automation, distributed applications, API integrations, and production infrastructure.
04

Applied LLM evaluation & research

Independent, reproducibility-focused evaluation work applying the same evidentiary standards used above to LLM behavior itself — the discipline behind the safety mechanisms shipped in VocalisAI.

Deployed LLM Eval Kit

2026
Open-source methodology
  • Built an open, reproducible methodology and toolkit for evaluating LLM safety behavior as deployed in product chat interfaces (not raw APIs): frozen rubrics, real blind labeling, and platform metadata captured as first-class results.
  • Implemented inter-rater concordance tooling (Cohen's kappa) and a stimulus-freezing workflow (SHA-256 hashing) to keep evaluation runs auditable and comparable over time.
github.com/zoharmx/deployed-llm-eval-kit

Behavioral Safety & Robustness Evaluation

2026
Independent Research
  • Designed a controlled cross-model evaluation of LLM safety behavior using frozen stimuli, repeated trials, captured outputs/metadata, and a pre-defined labeling rubric.
  • Built reproducible evaluation artifacts and explicit controls to distinguish model behavior from execution and prompt-scaffolding effects; established coordinated-disclosure procedures for security-relevant findings.

Deliberative LLM Evaluation

2026
Independent Research
  • Designed a falsifiable 25-case pilot comparing a multi-stage deliberative LLM architecture against direct model responses across five dilemma classes.
  • Implemented baseline comparison, operational risk-coverage metrics, component ablations, and signed, resumable checkpoints; audited the pipeline, identified parser-related execution artifacts, and recomputed results with documented provenance.

BLINDAJE — AI Agent Security Auditing

2026
Independent practice
  • Independent security-audit practice for deployed conversational agents: a 35-vector, 12-failure-family catalog covering prompt injection (direct and indirect), tool abuse, cross-user data leakage, and business-logic exploitation.
  • Every finding is evidence-linked to the exact conversation turn that produced it, with pass/fail criteria fixed before testing and a mandatory human-signed verdict; sample report available on request.
eduardorodriguez.site/blindaje
Operación BLINDAJE

This methodology, available as a service

BLINDAJE is the commercial arm of this research: a security audit for conversational AI agents already in production, run against a catalog of 35 attack vectors across 12 failure families. Pass/fail criteria are fixed before testing, every finding is linked to the exact conversation turn that produced it, and no report ships without a human-signed verdict.

05

Technical stack

Languages
Python·TypeScript·JavaScript·SQL
AI / LLM
OpenAI·Gemini·multi-LLM orchestration·AI agents·tool/function calling·structured outputs·LLM evaluation
Frameworks & Backend
FastAPI·asyncio·WebSockets·REST APIs
Observability / LLMOps
Langfuse·Datadog·structured logging·evaluation pipelines
Cloud / Infrastructure
Google Cloud (Cloud Run, Cloud Build)·Firebase/Firestore·Docker
Data
PostgreSQL·Supabase·Redis
Frontend
React·TypeScript·Next.js
DevOps
Git·GitHub·GitHub Actions·CI/CD
06

Recognition

Google Cloud AI + Datadog Hackathon — Finalist
Google Gemini Live Agent Challenge 2026 — Finalist

Looking for someone to own an agentic system end-to-end?

From asynchronous backend architecture to the evaluation harness that proves it works.