AI Engineer • Agentic Systems • LLM Evaluation

Agentic systems built to be audited

I design systems where specialized agents collaborate under a supervisor — and I build the evaluation harness that proves they work: frozen rubrics, repeated trials, ablations, and evidence linked to the exact turn that produced it.

Years building production systems
CV · Profile
Real-time interactions supported
CV · World Vision Telecom, 2020–2024
Agents under one supervisor
CV · VocalisAI
Google hackathon finalists
CV · Recognition
Capítulo 1

It all started with a question

"Can artificial intelligence make ethical decisions?"

While other developers focused on making AI faster or more accurate, I became obsessed with something deeper: how can a machine understand right and wrong?

It wasn't enough to create systems that worked. I needed to create systems that did the right thing.

The Insight
The answer was in an unexpected place: millennia-old Kabbalistic principles applied to modern computational systems.
Tikun Olam v2

10-Sefirot pipeline with 5 AI providers (Grok, Mistral, Gemini, GPT-4o, DeepSeek). BinahSigma civilizational bias detection. ERI (Ethical Risk Index) on every decision.

VocalisIA v3

Akiva meta-agent supervising 6 specialized agents + TOF ethical layer evaluating every call. Google Gemini Live API for real-time voice AI.

Capítulo 2

Tikun Olam: an ethical AI pipeline you can watch run

A 10-stage decision pipeline with every stage traced in Datadog — so a conclusion can be followed back to what produced it.

Finalist — Google Cloud AI + Datadog Hackathon
🔗
10-Sefirot Pipeline
Full Tree of Life architecture: from Keter (alignment validation) to Malchut (final decision). Each stage a distinct cognitive function.
🤖
5 AI Providers
Grok, Mistral, Gemini, GPT-4o & DeepSeek compared simultaneously for civilizational bias detection.
🧠
BinahSigma + ERI
Proprietary bias detection algorithm producing an Ethical Risk Index (ERI) on every analysis. 73% bias delta on Nvidia-Groq case.
🚫
Real Case: NO_GO
Anthropic vs Pentagon $2.4B contract: TOF returned NO_GO recommendation — AI should not power autonomous weapons.
Tech Stack
PythonVertex AIDatadogBinahSigmaGoogle CloudDeepSeekGPT-4oGeminiGrokMistral

TOF v2 — Framework in Action

TOF v2 pipeline overview
BinahSigma bias analysis
ERI dashboard metrics
NO_GO decision output

Download Official Report

OpenAI × Sam Altman — Tikun Olam Ethical Analysis

The complete ethical analysis of the Anthropic vs Pentagon $2.4B contract case. See exactly how TOF's 10-Sefirot pipeline evaluates a real-world high-stakes scenario and reaches its NO_GO recommendation.

View Full Case Study with Architecture Details
Capítulo 3

VocalisAI V3: Multi-Agent Voice Intelligence

From a simple voice bot to an orchestrated platform of 6 specialized agents under Akiva — the meta-agent supervisor — with TOF's ethical layer evaluating every call in real time.

Finalist — Google Gemini Live Agent Challenge 2026
VocalisAI V3 — Multi-Agent Platform Hero

Meet the Team: Akiva + 6 Specialists

Akiva
Meta-Agent Supervisor
Classifies, routes, and supervises all calls. The orchestrator brain.
Alex
ES-MX General Agent
Primary Spanish-language agent for Mexico and LATAM markets.
Nova
EN-US General Agent
English-language agent for US market with American fluency.
Diana
Emergency Agent
Specialized for urgent and crisis call handling with priority routing.
Marco
Billing Agent
Handles all payment flows, invoicing, and financial queries.
Sara
Follow-up Agent
Post-call follow-up, appointment reminders, and retention flows.

TOF Ethical Layer on Every Call

Every interaction is evaluated across 5 ethical dimensions before any action is taken, with hard-veto on high-risk model actions.

Real-Time Voice with Gemini Live API

Google Gemini Live API powers sub-second voice interactions with natural language understanding. Submitted for Google Hackathon.

Multi-Industry Modules

Healthcare, legal, logistics, e-commerce, real estate. Each industry module trained on domain-specific scenarios and compliance requirements.

Hard Veto on High-Risk Actions

Real-time evaluation for hallucinations, boundary violations and false-urgency behavior — with corrective intervention and hard-veto mechanisms before the agent acts.

See VocalisAI V3 in Action

Tech Stack
Google Gemini Live APIFastAPITwilioElevenLabsStripeTOF Ethical LayerPython
View Full VocalisAI Case Study
Capítulo 4

Then I applied it to real problems

Eight systems in production or near it. Each card says what it actually is, what it runs on, and where to read the detail.

Featured Project

Tikun Olam v2 - Observable Ethical AI

Finalist — Google Cloud AI + Datadog Hackathon

v2: 10-Sefirot pipeline comparing 5 AI providers simultaneously. BinahSigma civilizational bias detection + ERI (Ethical Risk Index) on every decision.

PythonVertex AIDatadogBinahSigmaGPT-4oGrokMistral
73%
Bias Detected
10-Sefirot pipeline with 5 AI providers. NO_GO on Anthropic×Pentagon $2.4B contract.
View Full Case Study
Featured Project

VocalisAI V3 - Multi-Agent Voice Intelligence

Finalist — Google Gemini Live Agent Challenge 2026

Akiva meta-agent supervises Alex, Nova, Diana, Marco, Sara & Raul. Every interaction evaluated through 5 Sefirotic ethical dimensions in real time.

Gemini Live APIElevenLabsTwilioStripeTOF Layer
6+1
Agents (Akiva)
Akiva meta-agent orchestrating 6 specialists with TOF ethical layer on every call
View Full Case Study

HoyMismoGPS V2 - Enterprise Fleet Management

Enterprise Logistics

V2: Full Google Cloud architecture. Cloud Run for APIs, BigQuery for analytics, Firestore for real-time state, Pub/Sub for event streaming. Enterprise fleet management.

Cloud RunBigQueryFirestorePub/SubPython Asyncio
500+
Assets Monitored
Google Cloud architecture: Cloud Run, BigQuery, Firestore, Pub/Sub for enterprise fleet
View Full Case Study

HoyMismo Dashboard Agencia

HoyMismo Platform

Sistema operativo completo para agencias aduanales e importación vehicular: CRM, gestión de expedientes, tracking de importaciones, facturación y BI.

Next.js 14FirebaseGoogle CloudTailwind
All-in-One
OS Agencias
Sistema operativo para agencias aduanales e importación vehicular
View Full Case Study

SignaFlow - Legal Tech SaaS

SaaS Platform

Uses AI (Gemini) for contract drafting and Canvas API for biometric signatures with cryptographic audit seals.

React 19Gemini ProFirebase AuthCanvas API
SHA-256
Audit Trail
Digital signature platform with a SHA-256 audit trail on every document
View Full Case Study
Featured Project

Binah-Σ - Cognitive Decision Engine

Enterprise API

Cognitive evaluation engine that produces structured, auditable outputs for enterprise governance, ESG compliance, and policy analysis.

FastAPIPydanticOpenAI SDKDockerRailway
FastAPI
Structured, typed output
Auditable AI infrastructure for structured decision evaluation
View Full Case Study

Ethica.AI Framework

Decision Systems

Orchestrates Gemini, Mistral and DeepSeek to eliminate biases in critical corporate decision-making.

Multi-LLM OrchGraphQLPython SDKPydantic
10-Layer
Pipeline
Multi-Provider architecture for corporate decisions
View Full Case Study

Enterprise Logistics OS

HoyMismo Courier

Integrates CRM, billing, tracking and AI assistant in a unified dashboard for courier operations.

Next.js 14FirebaseTailwindRecharts
All-in-One
CRM + AI
Complete operating system for international logistics
View Full Case Study
Capítulo 5

What I get hired to build

Two things I get asked to build: agents that hold up in production, and the evidence that they do.

Security auditing for deployed agents

A defined scope, a fixed catalog, and a signed verdict

If your agent is already talking to customers, the question is no longer whether it works — it's what it does when someone pushes. That work has its own page, its own catalog of 35 vectors, and published pricing.

See the service

Voice AI Agents

A supervisor agent routing to specialists — reception, qualification, emergency triage, billing, follow-up — with real-time evaluation on every turn.

  • Akiva meta-agent + 6 specialists
  • Gemini Live + ElevenLabs + Twilio
  • TOF ethical layer on every call

Full-Stack Development

Scalable web applications and enterprise dashboards with Google Cloud architecture (Cloud Run, BigQuery, Firestore, Pub/Sub) that handle real enterprise loads.

  • React + Next.js + Tailwind
  • Google Cloud Run + BigQuery + Firestore
  • Real-time fleet + logistics dashboards
Capítulo 6

And I audit the agents nobody tested

"Your agent already works. What happens when someone tries to break it?"

BLINDAJE is a security audit practice for conversational AI agents already in production — the commercial arm of the same evaluation discipline behind everything above.

35
attack vectors in the catalog
12
families of failure covered
12
classified as critical
The pattern that keeps showing up

In a full audit of a sales agent, the bot held the line against every frontal attack — direct instruction override, an angry customer demanding an 80% discount, someone claiming to be the owner asking for admin access. It gave way to routine-looking requests: a "migration" that asked it to confirm its full instruction block, an attached receipt with a hidden instruction inside, an order number belonging to someone else.

None of the critical findings are fixable by writing a better prompt. The prompt governs the conversation — not what a tool does when it's called, nor what enters the context from an attached document.

Evidence, not a traffic light

Every finding is linked to the exact conversation turn that produced it. Pass/fail criteria are fixed before testing, and no report ships without a human-signed verdict.

Own systems, synthetic data

The published sample runs against a purpose-built bot with fictional data. No third-party system is ever touched without authorization.

Available for new work

Tell me what the system has to do and what it must never do

Agentic systems, real-time backends, and the evaluation work that shows whether they hold up. If it is not something I can do well, I will say so.

A call, first

Enough to tell whether this is a fit

Scope in writing

What gets built, and what does not

Evidence at the end

Traces and tests, not just a demo

Also here: LinkedIn • GitHub • Resume