Javier Pontón González·Forward Deployed Engineer · AI Engineer
I build AI
that survives your audit
The evaluation harness exists before the first agent does, every extracted figure traces back to its document and page, and the models run air-gapped on your own hardware.
I embed with your team and build on your infrastructure: AI agents, RAG, MCP, evaluation, and model serving on Bedrock, Vertex AI or your own GPUs.
I am the only engineer on four live products, from the NFC access board I designed to the model that reads the documents.
- Seniority
- Senior / Lead · 7+ yrs backend, 2+ yrs production AI
- Role fit
- Forward Deployed Engineer · AI Engineer · LLM Engineer
- Availability
- Available immediately · 40 h/week
- Location
- Asturias, Spain · CET · fully remote (EU, UK, US)
- Right to work
- EU citizen · no sponsorship needed
- Engagement
- Independent Consultant (Freelance / B2B Contractor)
- Languages
- Spanish (native) · English (professional working proficiency)
- Notice
- None. Current engagement ends on agreement.
What the work produced.
Numbers from systems that went to production, not from slide decks. Each one belongs to a named engagement below.
Field-level scores on the golden dataset before go-live, air-gapped multi-agent tender platform
Invoice extraction in production with a self-hosted vision-language model
Azure to AWS migration of a global logistics platform
Consumer app shipped solo to both stores, three weeks after launch
Historical records moved from PostgreSQL to DynamoDB with the hot read path kept fast
Financial transactions through serverless batch pipelines on the T24 core
Garment-sorting events across the logistics network
Automated test coverage sustained under real TDD
Fibre installation platform running under international SLAs
Eight engagements, no gaps.
Independent Consultant (Freelance / B2B Contractor) since November 2023, back to back, contracted directly with no agency layer. Open any engagement.
2026 —IberiaForward Deployed Engineer (FDE) · fare pricing platform→
Technical owner of the fare pricing domain. Built and deployed an internal Model Context Protocol server exposing pricing tooling to LLM clients, enabling governed function calling over fares, providers and configuration instead of ad-hoc scripts.
Delivered a RAG service over internal specs with embeddings, hybrid search and cross-encoder reranking, instrumented with OpenTelemetry and Dynatrace like any other production service. Introduced Claude Code for scaffolding, Karate and JUnit generation and large refactors under conventions I defined for the team.
Deployed managed model serving on AWS Bedrock and Google Vertex AI to production alongside self-hosted inference, choosing per use case on latency, cost, data privacy, scalability and freedom of model choice, wired into the MCP server and the RAG service under the same OpenTelemetry and Dynatrace observability.
Owned resilience of the pricing path against third-party provider degradation: Kafka and PostgreSQL on AWS in hexagonal architecture with DDD, Redis caching on the highest-traffic flows.
2026 ∥Systems integrator · Saudi ArabiaForward Deployed Engineer · air-gapped multi-agent platform→
Architected a two-agent system for public-tender response: a financial agent performing structured extraction from vendor quotations across PDF, Word, Excel, email and supplier portals into a costing model; a technical agent applying RAG and Arabic OCR over RFPs of hundreds of pages.
Engineered source-level traceability from every extracted figure back to its document and page, mandatory human review gates, and visual flagging of AI content with no vendor source behind it. Evaluation harness defined before the build: golden dataset of historical tenders, field-level precision and recall, regression runs on every prompt or model change.
Benchmarked Qwen3, DeepSeek, Falcon-H1 Arabic and ALLaM, sized GPU and VRAM for vLLM and Ollama serving, and met Saudi PDPL and NCA requirements in a fully isolated network.
Sole engineer on the engagement, run in parallel with Iberia. The golden dataset of historical tenders scored 90%+ field-level precision and recall before go-live.
2025–26KPNSenior Full Stack Engineer · fibre installation platform→
Built semantic search with embeddings over runbooks and ServiceNow incident history, cutting the time to find the right precedent when triaging field-operations incidents. Introduced LLM-assisted development across the international team with shared prompt and review conventions, so AI output was always human-verified before merge.
Delivered a fibre installation tracking platform running in three European countries under international SLAs, and migrated historical production records from PostgreSQL to DynamoDB with chunk-based storage to keep the hot read path fast.
2024–25MercadonaLead Backend Engineer · product analytics→
Engineered a Spring Batch ingestion pipeline processing millions of product records per day from heterogeneous sources, with restartable jobs and idempotent writes. Designed API-first REST services in hexagonal architecture with DDD, sustaining over 85% automated test coverage under real TDD.
2023–24InditexLead Backend Engineer · global logistics→
Technical lead for the Azure to AWS migration, delivered with zero downtime. Engineered a serverless garment-sorting event platform processing thousands of events per minute across the logistics network, with async REST APIs in hexagonal architecture and Karate integration testing alongside the existing JUnit suite.
2022–23OpenbankSenior Backend Engineer · Santander Group→
Serverless batch pipelines on AWS Lambda processing millions of financial transactions per day on top of the T24 core banking system, inside a 100+ backend engineer organisation. Technical lead for the Java to Kotlin migration across services and Lambdas; delivered internal tech talks on DDD and hexagonal architecture to around 50 engineers.
2019–22Idealista · Empathy.coBackend Engineer · the foundation→
Idealista (2021–22). Digital contract-signing platform for Spain, Portugal and Italy: REST APIs plus Kafka event streams for real-time signing, with event sourcing keeping an auditable, replayable history of every contract.
Empathy.co (2019–21). Built the configuration platform enterprise clients including Kroger, Carrefour and Inditex use to tune their Elasticsearch-backed search engines, and executed zero-downtime migrations from Java 8 to 11, GCP to AWS and monolith to API Gateway.
Why engineering teams bring me in.
Forward deployed, not advisory
I work inside your team and on your infrastructure. Discovery with the people who own the workflow, scoping against your real data, and a handover your engineers can actually carry.
AI plus full-stack engineering
I don't treat AI as a separate prototype layer. I understand the backend, APIs, databases, distributed systems, infrastructure and production constraints around it.
Production over demos
Evaluation, observability, traceability, human review, reliability, latency and cost are part of the architecture, not a phase two.
End-to-end ownership
I take a system from architecture to implementation, deployment and production operation, including the incident at three in the morning.
I actually ship products
Beyond client work I build and operate my own products: real users, real payments, real App Store review, real GDPR obligations.
Proof, not promises.
A few systems I've actually built, shipped and still operate.
ZORRO
A consumer product, built and shipped solo.
Semantic matchmaking with embeddings and pgvector plus cross-encoder reranking: the same retrieve-then-rerank architecture as a production RAG pipeline, applied to people instead of documents. A self-hosted vision-language model moderates user photos, because sexual orientation is GDPR special-category data and never reaches a third-party processor.
Facturias
Turning invoices into validated data.
A self-hosted Qwen2.5-VL reads invoices as documents rather than OCR text dumps, preserving tables, line items and stamps. RAG over the Spanish tax code justifies every classification with a retrievable rule; every field carries a confidence score, and anything under threshold is routed to human review instead of written silently.
Apunta
From custom hardware to production software.
A multi-tenant platform for shooting clubs owned end to end: an NFC access PCB designed around an ESP32-C6 and a PN532 reader with Power over Ethernet, firmware in C on ESP-IDF, Spring Boot services with per-club tenancy, and a React Native app in both stores.
Grabia
AI that stays local.
WhisperX transcription and pyannote diarisation feed a LangGraph agent that produces structured summaries, decisions and action items, with RAG over the meeting archive. It runs on a private node (Ryzen AI MAX+ 395, 128 GB unified memory, ROCm) serving Qwen3-30B under Ollama and mistral.rs. No audio or transcript leaves the machine.
AI is easy to demo.
Engineering it for production isn't.
Most applied AI fails in the same four places: nobody defined what correct looks like, nobody can trace an output back to its source, nobody instrumented it so it can be debugged at three in the morning, and nobody asked whether the data was allowed to leave the building. I close all four before the first agent exists.
Golden dataset and scoring rubric before the build. Field-level precision and recall.
Every figure links back to its document and page, or it gets flagged.
OpenTelemetry traces per agent step, latency and token cost per span.
Timeouts, retries and fallbacks against third-party provider degradation.
Token and latency budgets treated as product requirements, not surprises.
Air-gapped deployments meeting GDPR, PDPL and NCA constraints.
Review gates where a wrong answer costs money, not a silent write.
Benchmarked, quantised, sized for real GPUs. Swappable by design.
What I actually work with.
Every term here is something I have shipped to production or operated. Pick a layer.
Who you would
be hiring.
I live in Asturias and work as an independent consultant, contracted directly with no agency between me and the client. Since February 2026 my title at Iberia has been Forward Deployed Engineer, and I held the same title on an air-gapped engagement for a Saudi systems integrator that ran in parallel with it. Seven years of production backend engineering for Iberia, KPN, Inditex and Openbank sit underneath that, two of them shipping Generative AI. What I build outside client hours says the rest: I designed the NFC access board for Apunta around an ESP32-C6, wrote its firmware in C, and shipped the app to both stores. I run my own inference node, 128 GB of unified memory serving Qwen3-30B, because some of the data I work with is not allowed to leave the machine. I am taking an official M.Sc. in Artificial Intelligence Research part-time alongside the consulting work. If you have a problem in that shape, write to me and tell me what is going wrong.
What people ask before the first call.
The same answers an AI assistant will find in llms.txt and agent.json, written once and served to both.
01What does Javier Pontón do?→
Javier Pontón González is a Forward Deployed Engineer and AI Engineer based in Asturias, Spain, working fully remotely for clients in the EU, UK and US. He embeds with the customer team, scopes the use case against their real data, builds on their infrastructure and stays through production and handover. The systems themselves are AI agents and multi-agent workflows with LangChain, LangGraph and MCP, Retrieval-Augmented Generation with embeddings and cross-encoder reranking, structured document extraction with vision-language models, and model serving on AWS Bedrock and Google Vertex AI or fully self-hosted and air-gapped. He has seven years of production backend engineering behind that, for Iberia, KPN, Mercadona, Inditex, Openbank and Idealista.
02What is a Forward Deployed Engineer, and where has Javier worked as one?→
A Forward Deployed Engineer embeds with the customer team instead of delivering from the outside: scoping the use case against the customer real data, building on the customer infrastructure, and staying through production operation and handover. Javier has held the title on two engagements, which ran in parallel. At Iberia, S.A. he has been Forward Deployed Engineer since February 2026, technical owner of the fare pricing domain and of the MCP server and RAG service built on top of it. For a confidential IT systems integrator in Saudi Arabia, from July to August 2026, he was the sole engineer on an air-gapped multi-agent tender-response platform delivered inside the customer isolated on-premise network under Saudi PDPL and NCA requirements.
03Does Javier work with managed model serving as well as self-hosted inference?→
Yes, and the choice is made per use case. He deployed managed model serving on AWS Bedrock and Google Vertex AI to production at Iberia alongside self-hosted inference, deciding between them on latency, cost, data privacy, scalability and freedom of model choice, all of it integrated with the MCP server and the RAG service under OpenTelemetry and Dynatrace. At the other end of the range he delivered a fully air-gapped platform in Saudi Arabia with no external API call at all, benchmarking Qwen3, DeepSeek, Falcon-H1 Arabic and ALLaM under vLLM and Ollama.
04Is Javier Pontón available for hire or contract work?→
Yes. He works as an Independent Consultant (Freelance / B2B Contractor) invoiced from Spain, is available immediately for up to 40 hours a week, and works fully remotely across EU, UK and US time zones. He takes both long-running forward-deployed engagements and shorter scoped work such as a two-week AI feasibility assessment. He can be reached at [email protected] or +34 623 920 307.
05What is Javier's experience with RAG and AI agents?→
He built a production Model Context Protocol server and a RAG service over engineering documentation at Iberia, using embeddings, hybrid search and cross-encoder reranking, instrumented with OpenTelemetry and Dynatrace. As sole engineer for a Saudi systems integrator he architected a two-agent system for Arabic public-tender response, combining structured extraction from vendor quotations with RAG and Arabic OCR over RFPs of hundreds of pages, scoring 90%+ field-level precision and recall on the golden dataset before go-live. He also ships RAG in his own products: Facturias retrieves over Spanish tax regulation, and Grabia retrieves over a local meeting archive.
06Does Javier work with on-premise or air-gapped LLM deployments?→
Yes, this is a core specialisation. He designed and delivered a fully air-gapped multi-agent platform for a Saudi systems integrator under PDPL and NCA requirements: model-agnostic architecture against OpenAI-compatible APIs, benchmarking of Qwen3, DeepSeek, Falcon-H1 Arabic and ALLaM, GPU and VRAM sizing for vLLM and Ollama serving, and no data leaving the client network. His own products run self-hosted vision-language models for the same reason, including GDPR special-category data in ZORRO.
07How does Javier evaluate LLM systems?→
Evaluation is defined before the build. A golden dataset representing real inputs, field-level precision and recall as the scoring rubric, regression runs on every prompt or model change, LLM-as-a-judge and human evaluation where the output is subjective, and hit rate and MRR for the retrieval layer. Everything is instrumented with OpenTelemetry so latency, token cost and failure modes are visible per agent step.
08What products has Javier built and shipped?→
Four live products as founder and sole engineer. Facturias (facturias.es) is a multimodal invoice-processing SaaS with a self-hosted vision-language model and RAG over Spanish tax regulation. ZORRO (somoszorro.com) is a dating app for the gay and queer community on iOS and Android with semantic matchmaking and self-hosted photo moderation. Apunta (apuntapp.com) is a multi-tenant SaaS for shooting clubs owned end to end from custom NFC hardware to the mobile app. Grabia is a fully local meeting-intelligence pipeline running on a private inference node.
09What is Javier's technical stack?→
Python, Java, Kotlin and Elixir on the language side. FastAPI and Spring Boot for services. LangChain, LangGraph and MCP for agents. PostgreSQL with pgvector, Elasticsearch, Redis, DynamoDB and MongoDB for data. Kafka for events. vLLM, Ollama and mistral.rs for self-hosted inference. AWS, Azure and GCP with Docker and Kubernetes. Hexagonal architecture, DDD and TDD as the default way of building.
10Where is Javier located and which languages does he speak?→
He is based in Asturias, Spain, in the Europe/Madrid time zone, and works fully remotely with clients across the EU, UK and US. He is a native Spanish speaker with professional working proficiency in English.
Start with
the hard part.
Write to me with the messy version: the workflow, the constraint, the thing that has to be right. If an LLM is the wrong tool I will say so, and there is a two-week feasibility assessment for exactly that question. I read every message myself and answer within a day. Available immediately, fully remote.
