mech.app

The mech.app newsletter

Agentic AI, minus the noise.

Get practical field notes on AI agents, automation, developer tools and security delivered to your inbox.

No spam. Unsubscribe anytime.

AI Agents

HEMA's HAL: How a 100-Year-Old Retailer Deployed MCP Without Exposing AWS Credentials to Clients

HEMA anchored MCP authentication in Microsoft Entra ID while keeping AWS credentials server-side, solving the credential distribution problem.

Source: aws.amazon.com
HEMA's HAL: How a 100-Year-Old Retailer Deployed MCP Without Exposing AWS Credentials to Clients

HEMA, a century-old Dutch retailer, built HAL (an internal AI assistant) on Amazon Bedrock AgentCore and Model Context Protocol. The interesting part is not the assistant itself. It is how HEMA solved the credential distribution problem that blocks most enterprise MCP deployments.

MCP servers expose tools (functions, data sources, prompts) to AI clients. In local development, you run the MCP server on your laptop and point Claude Desktop at it. In production, you need to authenticate users, enforce access policies, and keep cloud credentials out of client machines. HEMA’s architecture shows one way to do this without reimplementing MCP’s protocol.

The Credential Distribution Problem

MCP servers typically need cloud credentials to access data sources. If you distribute those credentials to every client machine, you lose audit trails, can’t revoke access granularly, and violate most enterprise security policies.

HEMA’s HAL needed to query internal systems (HR data, inventory, developer portals) without giving every employee an AWS access key. The solution was to anchor authentication in Microsoft Entra ID (formerly Azure AD) and run all MCP servers behind a gateway that holds the AWS credentials.

Architecture: Gateway as Credential Boundary

HAL’s architecture has three layers:

  • Client layer: Desktop app or web interface. Users authenticate with Microsoft Entra ID. No AWS credentials on the client.
  • Gateway layer: Runs on AWS infrastructure. Receives authenticated requests from clients, validates tokens, and forwards tool calls to MCP servers. Holds AWS credentials for Bedrock and data sources.
  • MCP server layer: Standard MCP servers exposing tools. They receive requests from the gateway, not directly from clients.

The gateway acts as a security boundary. It translates user identity (from Entra ID) into AWS IAM permissions and enforces tool access policies. MCP servers see only authorized requests.

Authentication Handoff

The authentication flow looks like this:

  1. User opens HAL client and authenticates with Microsoft Entra ID.
  2. Client receives an Entra ID token (JWT).
  3. Client sends a tool call request to the gateway, including the Entra ID token.
  4. Gateway validates the token against Entra ID.
  5. Gateway maps the user’s identity to an internal access policy.
  6. Gateway calls the appropriate MCP server using its own AWS credentials.
  7. MCP server queries the data source (using AWS credentials held by the gateway).
  8. Gateway returns the result to the client.

The client never sees AWS credentials. The MCP server never sees the user’s Entra ID token. The gateway is the only component that understands both identity systems.

Tool Access Policies

HEMA’s gateway enforces tool access policies without reimplementing MCP’s protocol. Each tool exposed by an MCP server has an associated policy that maps to Entra ID groups or roles.

Example policy structure:

Tool NameRequired Entra ID GroupAWS IAM Role UsedRate Limit
query_hr_datahr-teamhal-hr-reader10/min
search_inventoryall-employeeshal-inventory-reader30/min
deploy_serviceplatform-engineershal-deployer5/min

The gateway checks the user’s group membership before forwarding the tool call. If the user lacks the required group, the gateway returns an error without calling the MCP server.

This approach keeps policy logic out of the MCP servers themselves. The servers remain simple: they expose tools and execute them when called. The gateway handles authorization.

State Management and Observability

HAL uses Amazon Bedrock AgentCore for orchestration. AgentCore manages conversation state, tool call sequencing, and retry logic. The gateway is stateless: it validates each request independently.

Observability is anchored in CloudWatch. The gateway logs every tool call with:

  • User identity (Entra ID subject)
  • Tool name
  • Timestamp
  • Response status
  • Latency

This gives HEMA an audit trail for every action HAL takes on behalf of a user. If a tool call fails, the logs show whether the failure was authentication, authorization, or execution.

Deployment Shape

HEMA runs the gateway on AWS Lambda behind an Application Load Balancer. Each MCP server runs as a separate Lambda function. This keeps the deployment simple: no Kubernetes, no long-running processes, no connection pooling.

The client is a desktop app built with Electron. It handles Entra ID authentication using Microsoft’s MSAL library and communicates with the gateway over HTTPS.

The MCP servers are standard Python or Node.js functions. HEMA wrote some custom servers for internal data sources and uses community MCP servers (like the GitHub MCP server) for external integrations.

Failure Modes

The most common failure mode is token expiration. Entra ID tokens expire after an hour. If the client doesn’t refresh the token, the gateway rejects the request. The client handles this by refreshing tokens proactively and retrying failed requests once.

The second failure mode is rate limiting. If a user exceeds the rate limit for a tool, the gateway returns a 429 status. The client shows an error message and suggests waiting before retrying.

The third failure mode is MCP server timeout. If an MCP server takes longer than 30 seconds to respond, the gateway cancels the request and returns an error. This prevents slow data sources from blocking the entire system.

Code Snippet: Gateway Token Validation

Here’s a simplified version of the gateway’s token validation logic:

import jwt
from jwt import PyJWKClient
from functools import lru_cache

@lru_cache(maxsize=1)
def get_jwks_client():
    return PyJWKClient(
        "https://login.microsoftonline.com/{tenant_id}/discovery/v2.0/keys"
    )

def validate_token(token: str) -> dict:
    jwks_client = get_jwks_client()
    signing_key = jwks_client.get_signing_key_from_jwt(token)
    
    payload = jwt.decode(
        token,
        signing_key.key,
        algorithms=["RS256"],
        audience="api://hal-gateway",
        issuer=f"https://login.microsoftonline.com/{tenant_id}/v2.0"
    )
    
    return {
        "user_id": payload["oid"],
        "groups": payload.get("groups", []),
        "email": payload.get("email")
    }

def check_tool_access(user_groups: list, tool_name: str) -> bool:
    required_group = TOOL_POLICIES.get(tool_name, {}).get("required_group")
    if not required_group:
        return False
    return required_group in user_groups

The gateway caches the JWKS (JSON Web Key Set) to avoid fetching it on every request. Token validation happens on every request. Group membership checks happen after validation.

Trade-offs

This architecture trades latency for security. Every tool call goes through the gateway, adding 50-100ms of overhead. For HEMA, this is acceptable because most tool calls involve querying slow data sources anyway.

The gateway is a single point of failure. If it goes down, HAL stops working. HEMA mitigates this by running the gateway in multiple availability zones and using health checks to route traffic away from unhealthy instances.

The gateway is also a bottleneck. If HEMA scales HAL to thousands of users, the gateway will need to scale horizontally. Lambda handles this automatically, but the cost increases linearly with request volume.

Technical Verdict

Use this architecture if you need to deploy MCP in an enterprise environment with existing identity providers and strict credential policies. It works well when:

  • You already use Microsoft Entra ID (or Okta, Auth0, etc.) for authentication.
  • You can’t distribute AWS credentials to client machines.
  • You need audit trails for every tool call.
  • You want to enforce tool access policies centrally.

Avoid this architecture if:

  • You’re building a local development tool where users control their own credentials.
  • You need sub-50ms latency for tool calls.
  • You want to avoid the operational overhead of running a gateway.
  • Your MCP servers don’t need cloud credentials (e.g., they only expose local file system tools).

For most enterprises evaluating MCP, the gateway pattern is the only practical way to deploy it. HEMA’s implementation shows that it works at production scale without requiring changes to the MCP protocol itself.