Explanation
Architecture diagram
S — Situation
While leading a cloud architecture engagement at Deloitte for a major Financial Services (BFSI) client, one of the key challenges I worked on was enabling enterprise adoption of GenAI across multiple teams and use cases.
The demand was growing, but every new GenAI application was creating the same enterprise concerns — data security, governance, model access, observability and cost.
We also had enterprise knowledge distributed across sources such as SharePoint, Confluence, internal databases, file systems and Amazon S3.
The concern was that if every application team built its own RAG pipeline, model integration, security controls and monitoring, we'd end up with duplicated engineering effort and inconsistent governance.
So the architectural problem I was solving was: how do we accelerate GenAI adoption while keeping enterprise security, governance and economics centralized?
T — Task
As the Solutions Architect, my responsibility was to define the target architecture and reusable guardrails for an enterprise GenAI platform.
It needed to support RAG-based applications, integrate with enterprise identity and security controls, provide governed access to AI services, and give us GenAI-specific observability and cost visibility.
At the same time, application teams needed enough autonomy to independently build and deploy their business use cases.
Because the broader environment spanned AWS, Azure and GCP, I needed to balance portability and standardization against the operational benefits of AWS-managed services.
A — Action
My first decision was to treat this as a platform architecture rather than a single GenAI application.
The principle was simple: centralize enterprise controls and reusable capabilities, while decentralizing application development.
I therefore structured the architecture into three functional planes:
- Application Plane
- Data Plane
- AI & Model Plane
With Security & Governance and Observability & FinOps as cross-cutting capabilities.
1. Application Plane
For the application layer, I chose Amazon EKS because we had multiple containerized orchestration services, an existing Kubernetes operating model, and a requirement for some deployment portability across clouds. I accepted the additional operational overhead because those benefits outweighed the simplicity of a serverless approach. At the edge, users authenticate against Entra ID to receive an OAuth2 access token. Requests pass through AWS WAF to API Gateway, which validates the JWT token at the edge before forwarding traffic over an API Gateway VPC Link providing private connectivity to an internal Application Load Balancer in front of the EKS cluster. Within EKS, the centralized Orchestrator manages the core execution pipeline in 5 distinct steps:
- App Authorization: Validating tenant and user workspace permissions.
- Context Assembly: Fetching relevant authorized vector chunks, conversation history, and system instructions.
- Policy-Aware Prompt Construction: Assembling dynamic prompt templates with compliance guardrails.
- Model Routing: Cost/latency-based routing (e.g., simple queries to lightweight models, complex tasks to flagship models).
- Response Validation: Enforcing groundedness, safety, data-leakage checks, and citation verifications before returning the payload.
The important design decision here was separating application orchestration from the underlying data and model services, so application teams weren't tightly coupled to either.
2. Data Plane (3-Tier Storage Strategy)
To avoid database bottlenecks and single points of failure, I designed a specialized 3-tier Data Plane:
Ephemeral Cache (ElastiCache for Redis): Serves as a semantic query-response cache with its vector-search capability, to intercept repeated queries before hitting the vector database or LLMs.
Vector Retrieval Store (OpenSearch Serverless): Stores document chunks, vector embeddings, and document-level ACL security metadata. Amazon Bedrock Knowledge Bases handles asynchronous document ingestion, chunking, and embedding generation into OpenSearch.
Persistent State Store (Amazon Aurora PostgreSQL with RDS Proxy): Manages relational metadata (workspaces, system prompts, RBAC policies, routing rules) and unstructured chat session history using native JSONB. I placed RDS Proxy in front of Aurora to pool and manage connection spikes coming from scaling EKS orchestration pods.
An important security consideration was enforcing authorization before retrieved content entered the model context. We ensured only permitted content was passed to the model by evaluating user context directly at the query boundary.
3. AI & Model Plane
For the model layer, I used Amazon Bedrock as the governed access layer rather than allowing individual applications to integrate directly with different foundation models.
This gave the platform a consistent model-access pattern and allowed us to control approved models and usage centrally.
We enforced Amazon Bedrock Guardrails at the model boundary to provide centralized protection against prompt injection, PII leakage, and toxic content.
This reduced application-level coupling to individual models while keeping the platform in control of AI service access. Application teams interact with Bedrock through standardized platform APIs, decoupling business logic from underlying foundation models.
4. Security & Governance
I deliberately made security and governance cross-cutting platform capabilities rather than something every application team had to reinvent.
That included:
- Enterprise identity
- IAM
- Workload identity
- EKS Pod Identity: Eliminating static AWS credentials by granting EKS pods least-privilege IAM roles directly.
- RBAC
- Secrets management
- Encryption
- Network isolation: All inter-service traffic flows over private VPC Endpoints / AWS PrivateLink
- Audit logging
- Policy-based access
I also applied defense in depth.
Identity established who the user was, application authorization determined what they could access, retrieval controls constrained the context, and AI safety controls addressed model-level risks.
The key principle was that the model should never receive data simply because the application technically had access to it — the context still had to be authorized for the requesting user.
5. Multi-cloud
For multi-cloud, I deliberately avoided trying to make everything completely cloud-agnostic.
We standardized the parts where portability had real value:
- Kubernetes for application deployment
- Terraform for infrastructure provisioning
But where cloud-native managed services provided operational efficiency (such as Amazon Bedrock, OpenSearch Serverless, and Aurora Serverless), I leveraged native AWS capabilities. For example, on AWS, Amazon Bedrock provided the managed foundation-model layer.
So the principle was:
Portable where appropriate, cloud-native where valuable, rather than forcing everything into a lowest-common-denominator abstraction.
6. Observability & FinOps
Finally, I treated observability as more than infrastructure monitoring.
We needed visibility into:
- Application health: CPU/Memory utilization, EKS ingress latency, request volumes, and tracing.
- API and model latency
- Errors
- Token consumption
- Usage patterns
- Cost
- GenAI & RAG Quality: Retrieval quality, groundedness scores, citation accuracy, token consumption, and model latency.
- Database Performance & FinOps: Streaming Aurora PG Performance Insights and slow-query logs to CloudWatch, while tracking per-team token costs and model usage
That gave the platform team visibility across three dimensions:
- Reliability
- AI quality
- Economics
This was particularly important because a GenAI application can be technically healthy while still having poor retrieval quality, excessive token consumption or unexpectedly high inference costs.
We also used semantic caching as part of the cost-optimization strategy, and the broader FinOps work contributed to a significant reduction in LLM inference costs.
R — Result
The result was a reusable enterprise platform that transformed our operating model: application teams could focus on business use cases while the platform team provided standardized security, model access, RAG, observability and cost controls.
Speed to Value
We reduced new GenAI application onboarding time from months to less than two weeks, allowing 15 distinct application teams to deploy production use cases within the first quarter.
Cost & Latency
The ElastiCache semantic caching layer achieved a 25–30% cache-hit rate for common enterprise queries, directly cutting LLM token consumption and reducing average API response latency by 35%.
Security & Compliance
Data-leakage risks are reduced and validated by InfoSec, providing a governed platform that prevented unapproved GenAI deployments.
Connection Efficiency & Resilience
Implementing RDS Proxy in front of Aurora PostgreSQL reduced active database connections by 75% during peak traffic surges, while Aurora's multi-AZ distributed storage met strict BFSI recovery time objectives ($RTO < 30\text$, $RPO = 0$).
Key Architectural Lesson
The biggest architectural lesson I took from this was that enterprise GenAI isn't primarily an LLM-selection problem.
It's a:
- Data problem
- Identity problem
- Security problem
- Governance problem
- Platform problem
- Observability problem
- Economics problem
— with the LLM being one component of the overall architecture.
Read More at Enterprise GenAI Platform

![How to create AWS Account in 2 min!! AWS Developer Certification Course 2020 KAUSTUBH SHARMA [L-02]](/_next/image?url=https%3A%2F%2Fi.ytimg.com%2Fvi%2FLqbDn03Qy4M%2Fhqdefault.jpg&w=3840&q=75)



