I design systems where specialized agents collaborate under a supervisor — and I build the evaluation harness that proves they work: frozen rubrics, repeated trials, ablations, and evidence linked to the exact turn that produced it.
"Can artificial intelligence make ethical decisions?"
While other developers focused on making AI faster or more accurate, I became obsessed with something deeper: how can a machine understand right and wrong?
It wasn't enough to create systems that worked. I needed to create systems that did the right thing.
The Insight
The answer was in an unexpected place: millennia-old Kabbalistic principles applied to modern computational systems.
Tikun Olam v2
10-Sefirot pipeline with 5 AI providers (Grok, Mistral, Gemini, GPT-4o, DeepSeek). BinahSigma civilizational bias detection. ERI (Ethical Risk Index) on every decision.
VocalisIA v3
Akiva meta-agent supervising 6 specialized agents + TOF ethical layer evaluating every call. Google Gemini Live API for real-time voice AI.
Capítulo 2
Tikun Olam: an ethical AI pipeline you can watch run
A 10-stage decision pipeline with every stage traced in Datadog — so a conclusion can be followed back to what produced it.
Finalist — Google Cloud AI + Datadog Hackathon
🔗
10-Sefirot Pipeline
Full Tree of Life architecture: from Keter (alignment validation) to Malchut (final decision). Each stage a distinct cognitive function.
The complete ethical analysis of the Anthropic vs Pentagon $2.4B contract case. See exactly how TOF's 10-Sefirot pipeline evaluates a real-world high-stakes scenario and reaches its NO_GO recommendation.
From a simple voice bot to an orchestrated platform of 6 specialized agents under Akiva — the meta-agent supervisor — with TOF's ethical layer evaluating every call in real time.
Finalist — Google Gemini Live Agent Challenge 2026
Meet the Team: Akiva + 6 Specialists
Akiva
Meta-Agent Supervisor
Classifies, routes, and supervises all calls. The orchestrator brain.
Alex
ES-MX General Agent
Primary Spanish-language agent for Mexico and LATAM markets.
Nova
EN-US General Agent
English-language agent for US market with American fluency.
Diana
Emergency Agent
Specialized for urgent and crisis call handling with priority routing.
Marco
Billing Agent
Handles all payment flows, invoicing, and financial queries.
Sara
Follow-up Agent
Post-call follow-up, appointment reminders, and retention flows.
TOF Ethical Layer on Every Call
Every interaction is evaluated across 5 ethical dimensions before any action is taken, with hard-veto on high-risk model actions.
Real-Time Voice with Gemini Live API
Google Gemini Live API powers sub-second voice interactions with natural language understanding. Submitted for Google Hackathon.
Multi-Industry Modules
Healthcare, legal, logistics, e-commerce, real estate. Each industry module trained on domain-specific scenarios and compliance requirements.
Hard Veto on High-Risk Actions
Real-time evaluation for hallucinations, boundary violations and false-urgency behavior — with corrective intervention and hard-veto mechanisms before the agent acts.
See VocalisAI V3 in Action
Tech Stack
Google Gemini Live APIFastAPITwilioElevenLabsStripeTOF Ethical LayerPython
Eight systems in production or near it. Each card says what it actually is, what it runs on, and where to read the detail.
Featured Project
Tikun Olam v2 - Observable Ethical AI
Finalist — Google Cloud AI + Datadog Hackathon
v2: 10-Sefirot pipeline comparing 5 AI providers simultaneously. BinahSigma civilizational bias detection + ERI (Ethical Risk Index) on every decision.
PythonVertex AIDatadogBinahSigmaGPT-4oGrokMistral
73%
Bias Detected
10-Sefirot pipeline with 5 AI providers. NO_GO on Anthropic×Pentagon $2.4B contract.
V2: Full Google Cloud architecture. Cloud Run for APIs, BigQuery for analytics, Firestore for real-time state, Pub/Sub for event streaming. Enterprise fleet management.
Cloud RunBigQueryFirestorePub/SubPython Asyncio
500+
Assets Monitored
Google Cloud architecture: Cloud Run, BigQuery, Firestore, Pub/Sub for enterprise fleet
Two things I get asked to build: agents that hold up in production, and the evidence that they do.
Security auditing for deployed agents
A defined scope, a fixed catalog, and a signed verdict
If your agent is already talking to customers, the question is no longer whether it works — it's what it does when someone pushes. That work has its own page, its own catalog of 35 vectors, and published pricing.
A supervisor agent routing to specialists — reception, qualification, emergency triage, billing, follow-up — with real-time evaluation on every turn.
Akiva meta-agent + 6 specialists
Gemini Live + ElevenLabs + Twilio
TOF ethical layer on every call
Full-Stack Development
Scalable web applications and enterprise dashboards with Google Cloud architecture (Cloud Run, BigQuery, Firestore, Pub/Sub) that handle real enterprise loads.
React + Next.js + Tailwind
Google Cloud Run + BigQuery + Firestore
Real-time fleet + logistics dashboards
Capítulo 6
And I audit the agents nobody tested
"Your agent already works. What happens when someone tries to break it?"
BLINDAJE is a security audit practice for conversational AI agents already in production — the commercial arm of the same evaluation discipline behind everything above.
35
attack vectors in the catalog
12
families of failure covered
12
classified as critical
The pattern that keeps showing up
In a full audit of a sales agent, the bot held the line against every frontal attack — direct instruction override, an angry customer demanding an 80% discount, someone claiming to be the owner asking for admin access. It gave way to routine-looking requests: a "migration" that asked it to confirm its full instruction block, an attached receipt with a hidden instruction inside, an order number belonging to someone else.
None of the critical findings are fixable by writing a better prompt. The prompt governs the conversation — not what a tool does when it's called, nor what enters the context from an attached document.
Evidence, not a traffic light
Every finding is linked to the exact conversation turn that produced it. Pass/fail criteria are fixed before testing, and no report ships without a human-signed verdict.
Own systems, synthetic data
The published sample runs against a purpose-built bot with fictional data. No third-party system is ever touched without authorization.