AI data engineer. The guardrails, governance and observability around your model — delivered as code you can run, not a policy document you file.
Most AI projects stall in the same place: a prototype that impresses in a meeting but can't be trusted in production. No data guarantees, no guardrails, no observability, no audit trail.
And long before a regulator ever calls, something more immediate happens — an enterprise customer sends a security questionnaire, or an investor's technical due diligence asks how the model is monitored. That is where AI without a trust layer stops being a product and starts being a liability.
The bottleneck is no longer building AI. It's making it trustworthy. That's the gap I close — entirely in code.
Verified August 2026 · The deferral is Regulation (EU) 2026/1744, the Digital Omnibus on AI — in force 27 July 2026, amending Regulation (EU) 2024/1689. Penalties, accurately: the €35M-or-7% ceiling is Article 99(3), and it applies to the prohibited practices of Article 5. High-risk non-compliance is Article 99(4) — €15M or 3%. Both are whichever is higher, except for an SME, where Article 99(6) makes it whichever is lower. Anyone quoting 7% at a seed-stage company is selling fear, not compliance.
Not your product — the safety, governance, and reliability around it that let you ship AI you can answer for.
A risk register that names who could be harmed, and an override a person can actually reach.
Schema contracts and end-to-end lineage, so every decision traces back to data you can defend.
PII redaction, response grounding, and an eval harness that catches unsafe output before a user does.
Dashboards, drift thresholds and alerting — so degradation is caught before a customer reports it.
Access, audit trail and EU AI Act documentation generated from the system, not written beside it.
Credential-free suites, CI gates and full IaC. Behaviour is proven before it ships.
Idempotency, checkpointing and bounded automated remediation — recovery without a 3am page.
Seven dimensions, scored 0–100 — the same number the gauge moves from Demo to Production-ready. Every check is rated from absent to enforced-in-code, and every one maps to a real obligation under the EU AI Act, the GDPR, NIST AI RMF or ISO/IEC 42001. It is not a new standard; it is the existing ones, made testable.
A focused, fixed-scope review of your AI system against a responsible-AI production standard — risk and human oversight, data, guardrails, observability, governance, testing and reliability. You get a clear, prioritized picture of what stands between you and AI you can ship responsibly — and a plan to close the gaps.
End-to-end systems, all nine public on GitHub — code and CI open to inspection. More than 2,700 credential-free, CI-gated tests and full Infrastructure as Code across AWS, Azure, GCP, Databricks and Snowflake. Each one proves dimensions of the framework above.
An enterprise RAG compliance agent on AWS Bedrock, grounded in verbatim EUR-Lex regulation. A five-check deterministic gate stands between the model and the analyst: every cited article must exist in the retrieved context, and the agent may escalate but never soften a decision. 592 CI-gated tests.
Document intelligence for cross-border trade, on Textract, Bedrock and SageMaker. Every system like this eventually meets the same question — at what confidence do we publish? — and most answer it with a number that sounds safe. Here each field declares an error budget and the threshold is derived from it against a labelled set, with the upper bound and N printed beside it. Where no threshold fits the budget, the field is declared always-review — on this corpus, 31 of 36. Publishing everything is 26.72% wrong, a hand-picked 0.85 is 4.68%, derived is 0.22%. Deployed on real AWS in 49m 32s, 34/34 live checks, then destroyed. 527 tests, offline.
A multi-tenant regulated report factory on AWS Bedrock AgentCore — Runtime, Gateway, Identity, Memory, the Cedar policy engine and OTEL observability, all in Terraform. The model writes the sentence and marks the slot; deterministic code resolves every number through declared SQL over a pinned Iceberg snapshot, and the gate scans the rendered DOCX, not the data behind it. A contract may declare a lawful omission and never an internal failure — that is how “we could not compute it” quietly becomes “it was not material”. Deployed on real AWS in 21m 47s, 32/32 live checks, then destroyed. 489 tests, offline.
A real-time decision platform for an electricity distribution network — 250,000 meters, 2,000 EV chargers, 400 substations — on IoT Core, Kinesis and Managed Flink over an Iceberg lakehouse. One stream feeds three decisions that each need a different definition of “we have seen enough”: curtailment in seconds, meter anomaly in hours, settlement restated over days. Fail closed is the wrong reflex on a grid — a transformer keeps heating while nobody decides, so the safe state is a deterministic action that confesses to being one. 341 tests.
A LangGraph Supervisor/Architect/Infra/Medic state machine that designs, deploys and repairs pipelines on AWS, Azure, GCP and Databricks. When a deployment fails the Medic reads the real CI logs, patches the exact line and redeploys — and a fix is refused unless its evidence quote appears verbatim in real output. 380 tests.
One JSON contract drives Unity Catalog across three clouds and Snowflake, with a checker that proves both enforce the same capabilities. Governance is a gate, not a report: a credential-free analyzer fails any pull request that grants read on PII, shipped as the govgate CLI. 137 tests.
A Medallion lakehouse correlating vehicle telemetry with driver biometrics over a ±60-second temporal join, with a risk score that emits each factor’s contribution. GDPR Art. 9 is enforced, not documented — column masks NULL the biometrics, so the same query returns different data depending on who runs it. 173 tests.
A keyless, 100%-IaC streaming stack on GKE Autopilot: Kafka with Avro → Spark → Redis TimeSeries and BigQuery, with dbt marts. A z-test drift detector catches a silently miscalibrated sensor whose readings are all still in range. 102 tests.
An Airflow/PySpark ETL where a declared data contract is the single source of truth for validation, rejection lineage and PII pseudonymisation. Rejected rows aren’t dropped — they are quarantined with the rule they violated. 63 tests.
A logo tells you someone paid. It doesn't tell you what was built, or whether it holds up when someone checks. These are my own systems — which lets me do the thing a client engagement never allows: open the code, the CI runs and the failures, and let you check every claim yourself.
fleet_safety_officers. Every biometric comes back NULL and the location is
coarsened — while the risk score built from them is still returned. Nothing in the SQL changed;
only who ran it.
Read the repository →
Start with FintelliGuard: a compliance agent refusing its own model's output, live. No slides and no mock-ups — real terminals, real CI runs, real failures being caught. The other five are below, nine minutes in total.
Nothing downloads until you press play.
Not a list of things I have read about. Each of these appears in a public repository, in code you can open, under tests that run in CI.
Six certifications behind it: AWS GenAI Developer (Professional), AWS Data Engineer, Databricks GenAI Engineer, Databricks Data Engineer, HashiCorp Terraform, IAPP AI Governance Professional (AIGP).
MEng in Electrical & Computer Engineering from NTUA, and six certifications across AWS, Databricks, HashiCorp and the IAPP AI Governance Professional. My standards aren't a sales line — they're how every system above was built: security by default, observability from day one, tested before shipped, deterministic gates around probabilistic models, standards before speed.
I've also owned a budget, directed a national organisation's programs across every department, run its digital transformation, and presented to rooms of hundreds. That matters here for a specific reason: an audit is only useful if its findings survive the conversation with the people who have to fund them. Translating a business need into a delivered system is the part I started with, not the part I had to learn.