Responsible AI · In production · In code

Anyone can build an AI demo. I make it safe to ship.

AI data engineer. The guardrails, governance and observability around your model — delivered as code you can run, not a policy document you file.

All systems verified reference builds, public on GitHub 2,700+ CI-gated tests
your-ai-service Demo
41 % production ready
  • risk & human oversight
  • data quality & lineage
  • guardrails & safety
  • observability & drift
  • governance as code
  • tested in CI · infra as code
  • self-healing reliability
p99 latency 42 ms
Certified across AWS Databricks HashiCorp IAPP NTUA · MEng
The gap

The demo works. Production is another question.

Most AI projects stall in the same place: a prototype that impresses in a meeting but can't be trusted in production. No data guarantees, no guardrails, no observability, no audit trail.

And long before a regulator ever calls, something more immediate happens — an enterprise customer sends a security questionnaire, or an investor's technical due diligence asks how the model is monitored. That is where AI without a trust layer stops being a product and starts being a liability.

The bottleneck is no longer building AI. It's making it trustworthy. That's the gap I close — entirely in code.

Dec 2027
The AI Act's high-risk obligations were deferred — to 2 December 2027 for standalone systems, 2 August 2028 for AI embedded in regulated products. Sixteen extra months for standalone: the difference between building it properly and retrofitting it in a panic.

Verified August 2026 · The deferral is Regulation (EU) 2026/1744, the Digital Omnibus on AI — in force 27 July 2026, amending Regulation (EU) 2024/1689. Penalties, accurately: the €35M-or-7% ceiling is Article 99(3), and it applies to the prohibited practices of Article 5. High-risk non-compliance is Article 99(4) — €15M or 3%. Both are whichever is higher, except for an SME, where Article 99(6) makes it whichever is lower. Anyone quoting 7% at a seed-stage company is selling fear, not compliance.

Responsible AI, in code

I build the layer that makes AI responsible in production.

Not your product — the safety, governance, and reliability around it that let you ship AI you can answer for.

01

Risk & human oversight

A risk register that names who could be harmed, and an override a person can actually reach.

02

Data quality & lineage

Schema contracts and end-to-end lineage, so every decision traces back to data you can defend.

03

Guardrails & safety

PII redaction, response grounding, and an eval harness that catches unsafe output before a user does.

04

Observability & drift

Dashboards, drift thresholds and alerting — so degradation is caught before a customer reports it.

05

Governance as code

Access, audit trail and EU AI Act documentation generated from the system, not written beside it.

06

Tested & reproducible

Credential-free suites, CI gates and full IaC. Behaviour is proven before it ships.

07

Self-healing reliability

Idempotency, checkpointing and bounded automated remediation — recovery without a 3am page.

Where the trust layer sits Your AI system on the left feeds into the trust layer, which begins with risk and human oversight naming what could go wrong, and then the six controls that answer it — data contracts and lineage, guardrails, observability and drift, governance as code, tests and infrastructure as code, and automated recovery. That is what makes the output defensible to a customer, an enterprise security review, an investor's due diligence and a regulator on the right. Your AI system model · prompts · data the part you own THE TRUST LAYER — WHAT I BUILD 01 · RISK & HUMAN OVERSIGHT names the harm every control below is there to prevent Data contracts & lineage Guardrails Observability & drift Governance as code Tests & IaC Automated recovery Defensible to your customer to a security review to due diligence to a regulator IN THAT ORDER Where the trust layer sits Your AI system at the top feeds into the trust layer, which begins with risk and human oversight naming what could go wrong, and then the six controls that answer it — data contracts and lineage, guardrails, observability and drift, governance as code, tests and infrastructure as code, and automated recovery. That is what makes the output defensible to a customer, an enterprise security review, an investor's due diligence and a regulator at the bottom. Your AI system model · prompts · data the part you own THE TRUST LAYER — WHAT I BUILD 01 · RISK & HUMAN OVERSIGHT names the harm every control below is there to prevent Data contracts & lineage Guardrails Observability & drift Governance as code Tests & IaC Automated recovery Defensible to your customer to a security review to due diligence to a regulator IN THAT ORDER
The trust layer is not your product. It is the part that lets you answer for your product — and that last list — customer, security review, due diligence, regulator — is the order the questions actually arrive in.
The method

How I measure Responsible AI readiness.

Seven dimensions, scored 0–100 — the same number the gauge moves from Demo to Production-ready. Every check is rated from absent to enforced-in-code, and every one maps to a real obligation under the EU AI Act, the GDPR, NIST AI RMF or ISO/IEC 42001. It is not a new standard; it is the existing ones, made testable.

  • 01Risk & human oversight15%
  • 02Data quality & lineage15%
  • 03Guardrails & safety20%
  • 04Observability & drift15%
  • 05Governance as code15%
  • 06Tested & reproducible10%
  • 07Self-healing reliability10%
0 · Absent 1 · Ad-hoc 2 · Managed 3 · Production-grade
readiness-score
Demo0–40
Piloting41–70
Production-ready71–90
Audit-ready · EU AI Act91–100
Read the full framework See it scored on a real system
Try it

Score your own system in three minutes.

Fourteen questions — two per dimension, the two that separate a system you can defend from one you can only describe. You get a weighted score, a band, and your three weakest dimensions named.

It runs entirely in your browser. Nothing is sent anywhere — no account, no cookie, no request. Which is roughly the point.

your-system Unscored
% production ready

0 of 14 answered

A readiness report scoring Fleet Risk Lakehouse 68 out of 100, band Piloting: guardrails in red at 33.3, risk and governance amber, tested and reproducible at 100.
Not a mock-up. That is my own Fleet Risk Lakehouse, scanned by the tool on this site: 68/100, guardrails in red, the first two gaps named. A framework that cannot fail its author is a marketing device. Scan your repository →
Start here

Responsible AI Readiness Audit

A focused, fixed-scope review of your AI system against a responsible-AI production standard — risk and human oversight, data, guardrails, observability, governance, testing and reliability. You get a clear, prioritized picture of what stands between you and AI you can ship responsibly — and a plan to close the gaps.

  • ScopeRisk & human oversight, data, guardrails & safety, observability, governance, testing, reliability
  • You getA prioritized readiness report; the scored checklist, in a form you can re-score yourself in six months; a fix plan — what to change and in what order, with the first two weeks specified concretely enough that one of your engineers can start from it; and a started EU AI Act Annex IV technical file for your system
  • Timeline1–2 weeks, from the day read-only access lands
  • Price€8,000 excl. VAT — fixed, whatever the system turns out to need.
  • First threeThat is an introductory rate, and it is a trade rather than a discount: I publish the engagement as a case study — anonymised if you prefer, named if you're willing. Full scope either way. It holds for the first three engagements, through 31 December 2026.
  • ThenOptional implementation to close the gaps. The €8,000 is credited against implementation work of €20,000 or more, contracted within 60 days of the report.
Selected builds

Reference implementations, built to production standards.

End-to-end systems, all nine public on GitHub — code and CI open to inspection. More than 2,700 credential-free, CI-gated tests and full Infrastructure as Code across AWS, Azure, GCP, Databricks and Snowflake. Each one proves dimensions of the framework above.

80/80planted violations, every one refused by the named gate

FintelliGuard

An enterprise RAG compliance agent on AWS Bedrock, grounded in verbatim EUR-Lex regulation. A five-check deterministic gate stands between the model and the analyst: every cited article must exist in the retrieved context, and the agent may escalate but never soften a decision. 592 CI-gated tests.

AWS BedrockOpenSearch ServerlessMosaic AIKafka/MSKTerraform
FrameworkGuardrails & safety · Governance as code
Repo The 80 attacks Deep dive
5 · 31fields with a derived threshold · fields that publish nothing

Manifest

Document intelligence for cross-border trade, on Textract, Bedrock and SageMaker. Every system like this eventually meets the same question — at what confidence do we publish? — and most answer it with a number that sounds safe. Here each field declares an error budget and the threshold is derived from it against a labelled set, with the upper bound and N printed beside it. Where no threshold fits the budget, the field is declared always-review — on this corpus, 31 of 36. Publishing everything is 26.72% wrong, a hand-picked 0.85 is 4.68%, derived is 0.22%. Deployed on real AWS in 49m 32s, 34/34 live checks, then destroyed. 527 tests, offline.

TextractBedrockSageMakerStep FunctionsIceberg / AthenaRedshiftTerraform
FrameworkRisk & human oversight · Guardrails & safety
537numerals scanned in the rendered files · 0 untraceable

Attestor

A multi-tenant regulated report factory on AWS Bedrock AgentCore — Runtime, Gateway, Identity, Memory, the Cedar policy engine and OTEL observability, all in Terraform. The model writes the sentence and marks the slot; deterministic code resolves every number through declared SQL over a pinned Iceberg snapshot, and the gate scans the rendered DOCX, not the data behind it. A contract may declare a lawful omission and never an internal failure — that is how “we could not compute it” quietly becomes “it was not material”. Deployed on real AWS in 21m 47s, 32/32 live checks, then destroyed. 489 tests, offline.

Bedrock AgentCoreCedarIceberg / AthenaOpenSearch ServerlessTerraform
FrameworkData quality & lineage · Governance as code
0 of 3,779rows published before their interval closed · SQL, against the deployed table

Watermark

A real-time decision platform for an electricity distribution network — 250,000 meters, 2,000 EV chargers, 400 substations — on IoT Core, Kinesis and Managed Flink over an Iceberg lakehouse. One stream feeds three decisions that each need a different definition of “we have seen enough”: curtailment in seconds, meter anomaly in hours, settlement restated over days. Fail closed is the wrong reflex on a grid — a transformer keeps heating while nobody decides, so the safe state is a deterministic action that confesses to being one. 341 tests.

IoT Core / KinesisManaged FlinkIceberg / AthenaSageMakerLake FormationTerraform
FrameworkData quality & lineage · Risk & human oversight
4clouds, one code path

Self-Healing Multi-Cloud Agents

A LangGraph Supervisor/Architect/Infra/Medic state machine that designs, deploys and repairs pipelines on AWS, Azure, GCP and Databricks. When a deployment fails the Medic reads the real CI logs, patches the exact line and redeploys — and a fix is refused unless its evidence quote appears verbatim in real output. 380 tests.

LangGraphPinecone RAGTrinoKubernetesTerraform
FrameworkSelf-healing reliability · Observability & drift
Repo Eval report Deep dive
6/6crafted violations, every one refused · attacked on every run

Multi-Cloud Governance Platform

One JSON contract drives Unity Catalog across three clouds and Snowflake, with a checker that proves both enforce the same capabilities. Governance is a gate, not a report: a credential-free analyzer fails any pull request that grants read on PII, shipped as the govgate CLI. 137 tests.

Unity CatalogSnowflakeTerragruntOPA/RegoGenie
FrameworkGovernance as code · Guardrails & safety
Repo The analyzer Deep dive
NULLwhat an unprivileged principal gets back · GDPR Art. 9, in the query engine

Fleet Risk Lakehouse

A Medallion lakehouse correlating vehicle telemetry with driver biometrics over a ±60-second temporal join, with a risk score that emits each factor’s contribution. GDPR Art. 9 is enforced, not documented — column masks NULL the biometrics, so the same query returns different data depending on who runs it. 173 tests.

DatabricksSpark StreamingUnity CatalogManaged Grafana
FrameworkData quality & lineage · Governance as code
Repo The risk model Deep dive
0service-account keys, anywhere

Real-Time Telemetry Pipeline

A keyless, 100%-IaC streaming stack on GKE Autopilot: Kafka with Avro → Spark → Redis TimeSeries and BigQuery, with dbt marts. A z-test drift detector catches a silently miscalibrated sensor whose readings are all still in range. 102 tests.

KafkaSpark StreamingBigQueryRedis TimeSeriesGKE Autopilot
FrameworkData quality & lineage · Observability & drift
Repo The contract Deep dive
63credential-free CI tests across four jobs

Contract-Driven Data Pipeline

An Airflow/PySpark ETL where a declared data contract is the single source of truth for validation, rejection lineage and PII pseudonymisation. Rejected rows aren’t dropped — they are quarantined with the rule they violated. 63 tests.

AirflowPySparkdbtPostgreSQLGlue/Athena
FrameworkData quality & lineage · Tested & reproducible
Repo The contract Deep dive
All repositories on GitHub
The receipts

Three claims, and where to check them.

A logo tells you someone paid. It doesn't tell you what was built, or whether it holds up when someone checks. These are my own systems — which lets me do the thing a client engagement never allows: open the code, the CI runs and the failures, and let you check every claim yourself.

A GitHub pull request titled 'grant analysts read access to the CRM schema' with a failing required check
Governance is a gate, not a report. A pull request that grants read on a PII-classified schema turns the build red before review. Made required in branch protection, the red is a wall. Read the analyzer →
A Databricks SQL query selecting heart rate and stress score from the fleet live status table, run without membership of the fleet safety officers group: every biometric column comes back NULL while the risk score is still returned
GDPR Art. 9, enforced in the query engine. This is the result, not the configuration: the same query, run by a principal who is not in fleet_safety_officers. Every biometric comes back NULL and the location is coarsened — while the risk score built from them is still returned. Nothing in the SQL changed; only who ran it. Read the repository →
Per-field calibration table from Manifest: a threshold column of dashes, each row naming why no threshold could be derived, and a closing line reading 5 thresholds derived, 31 always-review
A number nobody could derive is not published. Every field declares an error budget; the threshold is the lowest score whose upper bound on published-and-wrong still fits it. Where none fits, the field publishes nothing — 31 of 36, the reason named per row. That is the column of dashes. Why it is derived →
Watch it run

Two minutes that show the whole idea.

Start with FintelliGuard: a compliance agent refusing its own model's output, live. No slides and no mock-ups — real terminals, real CI runs, real failures being caught. The other five are below, nine minutes in total.

Nothing downloads until you press play.

The stack

Everything below was used to build the nine systems above.

Not a list of things I have read about. Each of these appears in a public repository, in code you can open, under tests that run in CI.

GenAI & agents

  • AWS Bedrock Agents
  • Bedrock AgentCore
  • Bedrock Knowledge Bases
  • Bedrock Guardrails
  • action groups
  • MCP tool handlers
  • LangGraph
  • LangChain
  • LangSmith
  • Databricks Mosaic AI
  • Agent Framework
  • Genie
  • RAG & retrieval design
  • LLM eval harnesses
  • MLflow
  • XGBoost

Vector search

  • OpenSearch Serverless
  • Pinecone
  • Databricks Vector Search
  • Titan Text Embeddings v2
  • GTE-Large
  • chunking strategy
  • metadata filtering
  • grounding thresholds
  • citation provenance

Responsible AI

  • PII redaction
  • denied topics
  • prompt-attack filters
  • contextual grounding
  • prompt-injection mitigation
  • adversarial red-teaming in CI
  • drift: PSI · KS · z-test
  • EU AI Act Annex IV
  • EU AI Act Art. 12
  • model & dataset cards
  • GDPR Art. 9 · Art. 30
  • TreeSHAP

Data & streaming

  • Spark Structured Streaming
  • Kafka
  • Avro + Schema Registry
  • Airflow
  • dbt
  • Delta Lake
  • Apache Iceberg
  • Medallion architecture
  • Trino
  • Redis TimeSeries
  • BigQuery
  • Glue / Athena
  • data contracts
  • SCD-2
  • exactly-once

Cloud & governance

  • AWS
  • Azure
  • GCP
  • Databricks
  • Snowflake
  • Unity Catalog
  • column masks
  • row filters
  • OPA / Rego
  • Cedar
  • Lambda
  • MSK
  • ECS / EKS
  • AKS
  • GKE Autopilot
  • KMS · Secrets Manager

Infrastructure & CI

  • Terraform
  • Terragrunt
  • Databricks Asset Bundles
  • Docker
  • Kubernetes
  • GitHub Actions
  • keyless OIDC
  • Workload Identity Federation
  • pytest
  • offline eval harnesses
  • Grafana
  • Prometheus
  • Python
  • SQL
  • Bash

Six certifications behind it: AWS GenAI Developer (Professional), AWS Data Engineer, Databricks GenAI Engineer, Databricks Data Engineer, HashiCorp Terraform, IAPP AI Governance Professional (AIGP).

About

I work the way good production systems are built.

MEng in Electrical & Computer Engineering from NTUA, and six certifications across AWS, Databricks, HashiCorp and the IAPP AI Governance Professional. My standards aren't a sales line — they're how every system above was built: security by default, observability from day one, tested before shipped, deterministic gates around probabilistic models, standards before speed.

I've also owned a budget, directed a national organisation's programs across every department, run its digital transformation, and presented to rooms of hundreds. That matters here for a specific reason: an audit is only useful if its findings survive the conversation with the people who have to fund them. Translating a business need into a delivered system is the part I started with, not the part I had to learn.

Theofanis Tsakanikas
Theofanis TsakanikasAthens · remote, EU
Available now B2B engagements or employment Remote only · EU · CET/CEST Athens, Greece Greek native · English C2 · German C1
AWS GenAI Developer · Professional AWS Data Engineer Databricks GenAI Engineer Databricks Data Engineer HashiCorp Terraform IAPP AI Governance Professional