AI Agents Ate My Boilerplate: What Actually Works in Enterprise Banking

Five AI agents run in production at a MENA bank, after four years working with it remotely. Each is a Laravel queue worker, not an API endpoint, so models swap without a mobile release. App Store submission time fell from 14 days to four. Every decision carries an audit trail, and none act without a human.
11 min read · 2,112 words
I work remotely with an enterprise bank in the region, and have done for four years. I do not name it here. Where a detail would describe its security posture rather than the engineering, I have generalised it.
Three agentic AI vendors pitched us last quarter. Every pitch had the same shape. A fast demo, perfect context, and a price that assumed compute was free.
Then you ask about throughput. About audit trails. About what a hallucinated compliance decision costs at two in the morning while the transaction pipeline is already saturated.
In a regulated system the useful question is not what the agent can do. It is what the agent is allowed to decide on its own, and the honest answer is almost nothing.
What are the five agents, and what is each one allowed to decide?
Five agents run in production. None of them decide anything by themselves. Each produces a draft, a score or a suggestion, and a person signs it.
| Agent | What it produces | What it may never do |
|---|---|---|
| API contract enforcer | Laravel request classes, TypeScript types, breaking-change PRs | Merge anything touching a monetary amount |
| Compliance log archaeologist | Draft regulator reports with an evidence chain | File a report |
| Mobile release check | A submission confidence score and failing screenshots | Submit a build |
| Incident response shadow | Likely cause, suggested rollback, prior incidents | Execute a remediation |
| Code review agent | Findings on N+1 queries, authorization gaps, PII in logs | Approve a pull request |
That last column is the design. Everything else is implementation detail.
Continue Reading
How do you generate API contracts from documentation that lies?
By treating the documentation as a hint and the response samples as the truth. We integrate with more than 40 third-party APIs: card networks, local switches, government identity services. Their documentation includes PDFs scanned in 2019, Swagger files that disagree with the endpoint, and services that return HTTP 200 with an HTML error page in the body.
The agent runs inside a GitHub Actions workflow rather than a chat window. It pulls the current spec or scrapes whatever passes for documentation, generates strict TypeScript from real response samples, writes Laravel FormRequest classes with validation extracted from the field descriptions, and opens a pull request whose description is the breaking-change analysis.
It still invents required fields. Anything touching a monetary amount carries a HUMAN_VERIFY label and does not merge without a person reading it.
The win was not speed. It was that the mobile team stopped shipping type errors into production.
What does a regulator's audit request actually require?
Ninety days of transaction history, delivered inside 48 hours, with an evidence chain a person can follow back to the query that produced each number. That used to take four engineers three days. It now takes one agent and one compliance officer.
The agent has read-only views onto our transaction warehouse, ClickHouse underneath, and retrieval over our internal compliance playbook, vectorised in Postgres with pgvector. The output is a Markdown report with the SQL embedded, so the officer can rerun any figure rather than trust it.
Embedding the query next to the answer is the part worth copying. A number a reviewer cannot reproduce is a number that fails the audit, whatever produced it.
Why does data residency force inference onto your own hardware?
Because the transaction data cannot leave. That single constraint decides the model, the hardware and the budget, before anyone discusses quality.
So we serve an open-weight model on hardware we own, with vLLM in front of it. It is slower than a hosted frontier model and it costs more GPUs. It is also the only version of this that a regulator will accept.
Latency makes the same argument from a different direction. Regional fibre to Frankfurt runs around 60 ms on a good day, before the model has read a token. Hosted AI services are priced and designed around a latency budget that assumes you are next door to the datacentre. For anything a customer waits on, local inference is not a preference.
Data residency is not a compliance checkbox you satisfy at the end. It is the first architectural constraint, and it eliminates most of the vendor shortlist before the evaluation starts.
How do you stop shipping App Store rejections?
By testing against your own rejection history rather than against a generic checklist. Guideline 4.2, Minimum Functionality, is subjective, and it is applied unevenly across reviewers in our region. A rejection costs a launch window, not an afternoon.
The agent runs Maestro and Detox flows in CI. The judgement sits on top: it reads previous rejection reasons from the App Store Connect API, generates cases that target those specific patterns, flags screens that read as a wrapped web page, and produces a submission confidence score.
| Measure | Before | After |
|---|---|---|
| Submission to approval | 14 days | 4 days |
Apple did not get faster. We stopped submitting broken builds. The agent catches broken deep links, missing loading states and Arabic text truncating in right-to-left layouts. It cannot fix an app that is genuinely a wrapped website, and no agent can.
Should an agent remediate an incident on its own?
No. Not in a bank. Auto-remediation is a liability position, not an engineering one, so this agent is decision support for whoever is awake at three in the morning.
A PagerDuty webhook lands in a Laravel queue worker. The agent gathers recent deploys, database metrics and error logs, retrieves over our incident runbooks, and posts one Slack thread: likely root cause, a suggested rollback command, and the closest previous incidents.
During an annual bulk-payment run, card processing spiked 400%. The agent named the pattern correctly, bulk payment plus third-party processor throttling, and surfaced a manual failover procedure we had documented and never automated.
It also suggested a failover that would have broken our active-active architecture contract. A person caught it. That is the whole argument for keeping the last column of the table above empty.
What does a code review agent catch that generic tools miss?
The patterns that are only dangerous in your codebase. We tried the general-purpose review tools and they were too noisy to keep, because banking code has failure modes that a tool trained on public repositories does not recognise.
Ours takes PHPStan output as structured input and adds rules extracted from our own past incidents.
| What it catches | Why a generic tool misses it |
|---|---|
Eloquent relationships loaded in a loop without with() | It reads as valid code; only the query count shows it |
Authorization checks bypassed through DB::raw() | The gate exists, the query goes around it |
| Personally identifiable data in structured logs | A logging line looks harmless out of context |
| Transaction boundaries spanning an external API call | Correct locally, wrong once the network is involved |
Only the last two are really about banking. The first two are the same mistakes I write about in what yabasha.dev actually runs, at a scale where nobody gets fined for them.
What does "no code" actually cost by month six?
The pitch deck said no code. Here is the inventory six months in.
| What "no code" turned out to be | Why it exists |
|---|---|
| Around 400 lines of YAML | GitHub Actions orchestration for every agent |
| Terraform for the GPU cluster | vLLM does not provision itself |
| Laravel queue workers | Model calls time out, and a web request cannot wait |
| A custom evaluation framework | Every model swap has to be benchmarked before it ships |
"No code" is a demo conceit. A production agent needs infrastructure, observability and fallbacks, which are the same disciplines any distributed system needs.
The timeout row is the one that bit hardest, and I wrote up the general version of that failure in building production AI agents that actually work. A queued job whose timeout is shorter than the model call it makes does not fail loudly. It writes nothing.
Where do these agents still get it wrong?
In four places, consistently. I list them because the failures shaped the design more than the successes did.
| Failure | What it looks like | The guard |
|---|---|---|
| Invented required fields | A generated request class rejects valid traffic | HUMAN_VERIFY on anything monetary |
| Suggested remediation that breaks an architecture contract | A plausible failover that violates active-active | No agent executes anything |
| Schema mappings that are valid but business-wrong | Two fields that match on type and mean different things | Domain constraints as guardrails, not prompt text |
| Confidence without calibration | A high submission score on a build that still fails | The score advises, a person decides |
The third one is current work. I am testing whether agents can negotiate schema mappings across reconciliation vendors better than people do, on the theory that agents do not get bored and do not assume an equivalence is obvious. They converge fast. They also converge on mappings that pass every type check and are wrong about the business.
Why is every agent a queue worker rather than an API endpoint?
Because it lets us change the model without shipping a mobile release. Every agent in this stack is Laravel-side. The React Native app and the Next.js frontend call ordinary APIs, and the model is an implementation detail behind them rather than part of the interface.
| Agent as queue worker | Agent as API endpoint |
|---|---|
| Swap models without an app release | Model choice is baked into the client contract |
| Add a rule-based fallback or a human step behind the same API | Fallback logic has to live in the client |
| Scale GPU spend separately from user-facing compute | The two are coupled |
| Retries and timeouts are the queue's problem | Every caller reimplements them |
This is the same shape I use at a far smaller scale on my own site, described in how I built an AI agent for my portfolio, and the cost argument for it is in cutting LLM API costs 60%.
Nobody using the app cares whether a model wrote the response. They care whether their salary arrived. The architecture follows from that, not from the model.
What is different about building this in MENA?
Three constraints that changed the architecture rather than the roadmap.
| Constraint | What it forced |
|---|---|
| Connectivity that shifts under you | Retry logic that would look paranoid in Silicon Valley: routing changes, filtering updates, seasonal traffic patterns that break a statistical model |
| Several regulators, different guidance | SAMA, the Central Bank of Jordan and CBUAE each publish their own AI governance direction, so we design for evidence extraction rather than accuracy alone |
| We cannot hire fifty ML engineers | The stack optimises for Laravel developers who can read a model card; the agent infrastructure is deliberately boring |
The third constraint is the one that decided the most. A design only your two best people can operate is not a production design.
What would I check before trusting an agent in a regulated system?
Work from the audit backwards. If a reviewer cannot reproduce the number, nothing else on this list matters.
- Every agent output a person signs is reproducible from the query or evidence embedded beside it.
- No agent executes anything. It drafts, scores or suggests, and a person acts.
- Anything touching money carries an explicit human-verification gate.
- The data-residency constraint is settled before the model shortlist, not after.
- Every model call runs behind a queue, so timeouts and retries belong to the queue rather than the caller.
- Model swaps are benchmarked against your own evaluation set before they reach production.
- The failure list is written down, and each entry has a named guard.
Trust but verify is the old banking phrase. For agents I would add: automate, but audit. Every decision an agent makes without a person seeing it is a decision you cannot defend.
Measured with
Figures as stated in the original post, written from four years working remotely with an enterprise bank in the region. The bank is not named, and operational detail that would describe its security posture has been generalised.
- Laravel backend, Next.js web, React Native mobile
- More than 40 third-party API integrations
- Open-weight model served with vLLM on owned hardware; ClickHouse warehouse; Postgres with pgvector
- Regulator audit shape: 90 days of history inside 48 hours
- App Store submission to approval: 14 days before, 4 days after
- Six months of production agents at the time of writing
Questions: @bayyash

Bashar Ayyash (Yabasha)
AI Systems Architect for regulated industries — evals, harness design, AI security.
Bashar Ayyash is an AI engineer and dev lead in Amman, Jordan. 20 years shipping software, 4 years inside a regulated bank building production RAG and agent systems with evals, guardrails and monitoring — in Arabic and English. He writes at yabasha.dev and builds open-source tooling for AI-assisted development.
Newsletter
Practical AI + full-stack insights for MENA builders. No spam.
Related Articles

Your Call Logs Are the Only Training Data That Actually Matters

Building Production AI Agents That Actually Work

How I Built an AI Agent for my Portfolio (Yabasha.dev) using Laravel & Next.js

Graph Engineering Is Mostly Airflow With A New Coat Of Paint
Read more on the blog
Browse the latest articles or explore the full archive.