3.16

View in English

3.16 API gateways and service mesh

Overview and motivation

The moment you split one program into many services, a new question appears: who is in charge of the traffic between them, and the traffic coming in from outside? You can answer that question badly by scattering the same concerns (authentication, retries, timeouts, rate limits, logging) into every service by hand, or you can answer it well by pushing those concerns into a shared layer that every service inherits for free. This chapter is about two such layers. An API gateway sits at the front door and manages traffic coming in from clients. A service mesh sits between your services and manages traffic flowing among them. They solve related problems in different places, and confusing the two is a common and expensive mistake.

The industry names the two directions of traffic with a compass metaphor. North-south traffic is the traffic that crosses the boundary of your system: a mobile app, a browser, or a partner calling in. East-west traffic is the traffic that stays inside your system: service A calling service B calling service C to satisfy one request. An API gateway is the specialist for north-south. A service mesh is the specialist for east-west. Keeping that distinction sharp is the single most useful idea in this chapter, because it tells you which tool owns which policy, and it stops you from doing the same work twice.

For large teams, these layers are how you enforce a policy once instead of a hundred times. When authentication, encryption in transit, and rate limiting live in a shared layer, a security fix ships to every service the day you deploy the layer, rather than waiting on a hundred backlogs. In enterprise and government settings, that centralisation is often the point: auditors want a single, provable place where access is checked and traffic is encrypted, and a shared gateway or mesh gives them exactly that policy enforcement point. This chapter builds on the architectural styles of chapter 3.2, the distributed-systems realities of chapter 3.3, and the networking fundamentals of chapter 3.13, and it turns them into concrete guidance about who handles your traffic.

Key principles

  • Separate north-south (gateway) from east-west (mesh); let each own its direction.
  • Push cross-cutting concerns into a shared layer so you write them once, not per service.
  • Adopt a service mesh only when the number of services makes per-service wiring the bigger cost.
  • Define one policy enforcement point per concern; never let gateway and mesh both do the same job.
  • Keep services thin: the platform handles transport, the service handles business logic.
  • Prefer identity-based, zero-trust networking over trust based on network location.
  • Buy the operational complexity of a mesh with your eyes open, and measure whether it pays.

Recommendations

Understand what an API gateway does

An API gateway is a single entry point that sits in front of your services and mediates every request from the outside world. At its simplest it is a smart reverse proxy (a server that receives client requests and forwards them to the right backend), but a gateway earns its name by doing far more than forwarding. It routes each request to the correct service based on the path, host, or headers. It authenticates the caller (verifying who they are) and authorises the request (checking what they may do), so a service behind it can trust that a request already passed the front door. It enforces rate limiting (capping requests per client over time) and quotas (capping total usage over a longer window) so one noisy or abusive client cannot starve the rest.

A gateway also reshapes traffic. Request transformation rewrites headers, translates between protocols, or adapts an old client’s format to a new service’s expectations. API composition lets the gateway fan a single inbound request out to several services and stitch their responses into one, so a client makes one call instead of six. Versioning support lets you run v1 and v2 of an API side by side and route each client to the version it expects, which buys you room to evolve without breaking anyone. Concentrating these concerns at the edge keeps your services focused on business logic and gives you one place to observe, secure, and throttle everything that enters. The design of the APIs the gateway fronts is the subject of chapter 2.3, and the identity checks it performs draw on chapter 4.7.

Use the backend-for-frontend pattern for divergent clients

A single general-purpose API often serves a web app, a mobile app, and partner integrations at once, and it serves all of them slightly badly. The mobile client wants small payloads and few round trips because bandwidth and battery are scarce; the web client can handle chattier, richer responses; the partner wants a stable contract that never surprises them. The backend for frontend (BFF) pattern resolves this tension by giving each class of client its own thin gateway, tailored to that client’s needs, sitting in front of the shared services behind it.

A BFF is a gateway with a narrower audience. The mobile BFF composes and trims responses so the app makes one efficient call; the web BFF exposes a fuller shape; the partner BFF holds a slow-moving, carefully versioned contract. Each team can evolve its own BFF without waiting on the others, which is often the real win, because it decouples client teams from one another. The cost is more moving parts and some duplicated logic across BFFs, so reserve the pattern for cases where client needs genuinely diverge. When every client wants the same thing, one gateway is simpler and better.

Understand what a service mesh does

A service mesh manages the east-west traffic between your services, and it does so without asking those services to change their code. The classic mesh works by deploying a sidecar proxy (a small proxy process that runs alongside each service instance and intercepts all of its network traffic). Your service thinks it is talking directly to another service; in reality it talks to its local sidecar, which handles the real network call. Because every request now flows through a proxy the platform controls, the mesh can enforce behaviour uniformly across every service, in every language, with no shared library to keep in sync.

What does it enforce? First, mutual TLS (mTLS), where both sides of every connection present certificates and encrypt the traffic, so service-to-service calls are authenticated and private by default. Second, traffic management: the mesh can shift a small percentage of traffic to a new version for a canary release, split traffic by header for testing, or mirror traffic to a shadow service. Third, resilience: retries, timeouts, and circuit breaking (the patterns of chapter 2.20) applied at the platform layer, configured by policy rather than coded into each service. Fourth, observability: because every request passes through a proxy, the mesh emits consistent metrics, logs, and distributed traces for all service-to-service traffic, feeding the observability practices of chapter 9.2. The service author writes none of this and gets all of it.

Know the sidecar pattern and the sidecarless alternatives

The sidecar model is elegant but not free. Each service instance now runs an extra proxy container that consumes memory and CPU, and every call makes two extra network hops (into the local sidecar and out of the remote one), adding a little latency. At a handful of services this overhead is invisible; across thousands of pods it becomes a real line item in your compute bill and your latency budget. That cost has driven a wave of sidecarless, or proxyless, approaches.

Two directions matter. One moves mesh functions out of a per-pod sidecar and into a per-node proxy, so many services on the same machine share one proxy instead of each running their own; this trades some isolation for a large drop in overhead. The other, proxyless approach embeds the mesh’s logic directly into the service through a thin library or the runtime, removing the extra hops entirely at the cost of a per-language dependency. A newer development pushes some mesh functions into the operating system kernel using eBPF (a technology for running sandboxed programs inside the Linux kernel), which can enforce policy and collect telemetry with less overhead than a userspace proxy. You do not need to bet on one winner today. You do need to know that the sidecar tax is real, that alternatives exist, and that your platform choices should not lock you out of adopting them later. These patterns sit on the container orchestration foundation of chapter 8.3.

Decide when a mesh earns its complexity

A service mesh is powerful and genuinely complicated to run. It adds a control plane to operate, proxies to upgrade, certificates to rotate, and a new layer to debug when a request goes missing. That complexity is worth buying when you have enough services that wiring these concerns by hand, per service and per language, costs more than operating the mesh. The rough signal is scale and polyglot diversity: dozens or hundreds of services, written in several languages, where a shared library for mTLS and retries would be a nightmare to keep consistent. At that scale, a mesh pays for itself in uniformity and provable security.

The mesh does not earn its complexity when you have a handful of services, a single language, or a small team. For a modest system, a good library or framework can give you mTLS, retries, and metrics with far less operational burden than a full mesh, and a plain gateway plus sensible client libraries often covers everything you need. Adopting a mesh because it is fashionable, before your scale demands it, is a common way to spend a year operating infrastructure that solves a problem you do not have. Start with the gateway, add resilience patterns in code or libraries, and reach for a mesh when the count of services and languages makes the per-service approach the more expensive one. This is the same “does it earn its complexity” discipline the cloud and distributed-systems chapters (3.11 and 3.3) return to again and again.

Avoid double-handling where gateway and mesh overlap

Gateways and meshes overlap, and the overlap is where teams hurt themselves. Both can do retries, both can enforce timeouts, both can check identity, both can collect telemetry. If the gateway retries a request three times and the mesh also retries it three times at every internal hop, one client retry can explode into dozens of backend calls and turn a small blip into a retry storm. If both layers enforce a timeout and the inner one is longer than the outer one, the outer gives up while the inner keeps working, wasting effort on a response no one will read.

The fix is a clear division of labour written down and agreed. Assign each concern to exactly one layer. The gateway owns north-south concerns: end-user authentication, external rate limits and quotas, request transformation, and API composition for clients. The mesh owns east-west concerns: service-to-service mTLS, internal retries and circuit breaking, and traffic shifting between service versions. Where a concern could live in either, pick one owner and make the other layer pass through. Configure retry budgets and timeout hierarchies so an outer timeout is always longer than the inner work it waits on. The goal is that every request has exactly one place that handles each concern, and no request is retried, authenticated, or logged twice by accident.

Treat gateway and mesh as policy enforcement points for zero trust

The deepest reason to run these layers is security architecture. Zero trust is the principle that no request is trusted because of where it came from; every request must prove its identity and its authorisation, even inside your own network. The old model trusted anything already inside the perimeter, which meant one breached service could roam freely. Zero trust replaces network-location trust with identity-based trust on every hop, and gateways and meshes are the natural enforcement points where that identity is checked.

The gateway is the policy enforcement point for external identity: it verifies the end user or partner before anything reaches your services. The mesh is the policy enforcement point for workload identity: every service gets a cryptographic identity, mTLS proves it on every call, and policy decides which services may talk to which. Together they give you defence in depth, where a request is checked at the edge and again between services, so a compromise of one service does not grant free movement to the rest. In multi-cluster and multi-region deployments, a mesh can extend this identity fabric across cluster boundaries, so a service in one cluster authenticates to a service in another with the same mTLS guarantees it uses locally, giving you consistent zero-trust networking even as the footprint spreads. The identity foundations here connect directly to chapter 4.7.

Trade-offs: pros and cons

ApproachProsCons
API gatewayOne place for authn, rate limits, composition, versioningA single choke point to scale and keep highly available
Backend for frontendEach client gets a tailored, independently evolving APIMore gateways to run; logic duplicated across BFFs
Service mesh (sidecar)Uniform mTLS, retries, observability with no code changeProxy overhead, latency, and a control plane to operate
Sidecarless / proxyless meshLower overhead and latency than per-pod sidecarsLess mature; weaker isolation or per-language dependency
Library-based resilienceSimple to run; no extra infrastructurePer-language duplication; hard to keep consistent at scale
Mesh across clustersConsistent zero-trust identity everywhereSignificant operational and networking complexity

The central tension is uniformity versus operational cost. A shared layer buys you consistency, provable security, and policy you write once, but it is a real system you must run, scale, secure, and debug, and it inserts itself into the path of every request. Resolve the tension by scale and need. A gateway almost always pays off the moment you have external clients, because the concerns it centralises are ones you cannot avoid. A mesh pays off later, when the number of services and languages makes per-service wiring the more expensive path. Below that threshold, libraries and a plain gateway give you most of the benefit at a fraction of the cost. Above it, the mesh’s uniformity is worth its weight. The mistake in both directions is adopting on fashion rather than need: a mesh too early is a year of yak-shaving, and a mesh skipped too late is a hundred inconsistent, hand-rolled implementations of mTLS.

Questions to discuss with your team

  1. For each cross-cutting concern (authentication, retries, timeouts, rate limiting, encryption, telemetry), which single layer owns it, and can everyone name the owner without guessing? This is the question that prevents double-handling, and most teams have never answered it explicitly, which means the answer differs by service and by author. Bring a concrete list of concerns down the left side and your layers (client library, gateway, mesh, individual service) across the top, and fill in the grid together. The places where two cells are checked for one concern are your retry storms and timeout inversions waiting to happen. The empty rows are the concerns no one is handling at all. The deliverable is a single agreed table, published where every team can see it, that says exactly one layer owns each concern and the others pass through. That table is worth more than any amount of gateway or mesh configuration, because it is the thing that keeps the two layers from fighting each other.

  2. Do we actually have enough services and language diversity to justify a service mesh, or are we about to buy a control plane to solve a problem we do not have? A mesh is a serious operational commitment, and the honest answer for many teams is that a good library plus a gateway would serve them better today. Bring the real count of your services, the number of languages they are written in, and an honest assessment of how much duplicated networking logic actually hurts you right now. Then bring the other side: who will operate the mesh, upgrade its proxies, rotate its certificates, and be paged when it misroutes a request. If the pain of per-service wiring is smaller than the cost of running the mesh, you have your answer, and it is to wait. If you are drowning in inconsistent mTLS and retry code across dozens of polyglot services, the mesh earns its keep. The point is to decide on evidence, not on what a conference talk made the mesh look like.

  3. Where is our zero-trust boundary today, and what happens if one internal service is compromised? Many systems still trust anything already inside the network, which means a single breached service can move laterally and reach everything else, and teams often discover this only during an incident. Walk through the blast radius honestly: if an attacker owns one of your services, what can it call, what can it read, and what stops it? Bring your current answer to how service-to-service calls are authenticated and encrypted, and be specific about which calls are protected by mTLS and which are plaintext trust based on being on the same network. The action that follows is a deliberate plan for identity-based access on every hop, with the gateway checking external identity and the mesh or an equivalent checking workload identity, so that a compromise is contained rather than catastrophic. Even if you are not ready to run a full mesh, naming where the trust boundary really sits is the first honest step.

  4. If our API gateway tier failed right now, how much of the system goes dark, and have we tested that failure rather than assumed it away? The gateway centralises so much that its outage takes everything behind it offline, and large teams tend to under-invest in its redundancy precisely because it works quietly until it does not. Weigh the competing pulls: a single simple gateway is easy to reason about and cheap to run, while a horizontally scaled, multi-zone tier costs more and adds its own failover configuration and complexity. Bring real numbers to the discussion, including how many instances run today, across how many availability zones, what the failover time is, and when you last ran a game day that killed the gateway on purpose. For enterprise and government platforms carrying uptime commitments or statutory service levels, add the contractual or regulatory penalty for an outage, because a front door with no redundancy is an availability risk you have quietly accepted on behalf of every service and every citizen behind it.

  5. Have we measured the sidecar tax our mesh actually charges, and do we have a plan for the sidecarless and eBPF alternatives, or are we paying it blind? At scale the per-pod proxy overhead in compute and latency stops being invisible and becomes a real line in the budget, yet many teams run thousands of sidecars without ever measuring the cost. The tension is between the sidecar model’s maturity and isolation on one side and the lower overhead of per-node proxies, proxyless libraries, or kernel eBPF approaches on the other, which are newer and trade away some isolation or add a per-language dependency. Bring the measured figures: memory and CPU consumed by proxies across the fleet, the added tail latency per hop, and what fraction of your compute bill the mesh represents. For a large enterprise running the mesh across thousands of pods, or a government platform under budget scrutiny, this is a spending decision oversight will eventually ask about, so knowing the tax and whether your platform choices keep the cheaper alternatives open is basic due diligence.

  6. Does each client class genuinely need its own backend for frontend, or are we about to duplicate logic across gateways for clients that actually want the same thing? The BFF pattern decouples client teams and lets each evolve a tailored contract, but every new BFF is another gateway to run, secure, and keep in sync, and the duplicated logic across them quietly becomes a maintenance tax. The competing considerations are team autonomy and client-specific efficiency versus the operational cost and drift of many near-identical gateways. Bring evidence of how much the clients’ needs really diverge: payload sizes, round-trip counts, versioning cadence, and how often one client’s change would have blocked another under a single shared gateway. In a large organisation with many client teams, or a government platform serving a public web app, a mobile app, and partner integrations at once, the honest question is whether the divergence justifies the proliferation, because a BFF per client that all want the same shape is sprawl you will pay to maintain for years.

Sector lens

Startup. Ship an API gateway and skip the mesh. With a handful of services and a tiny team, a single gateway handles authentication, rate limits, and composition, while a shared client library gives you mTLS and retries for a fraction of a control plane’s cost. Your scarcest resource is engineering attention, so a mesh you cannot operate is a liability, not a moat. Keep the gateway redundant enough to survive one node dying, and revisit the mesh only when service count and language diversity actually force the question.

Small business. You have no platform team to run a mesh, so lean on what your hosting platform or gateway product gives you out of the box: managed TLS, built-in rate limiting, and a hosted gateway rather than one you patch yourself. Frame the choice as buy versus build, and buy, because a managed API gateway costs less than the engineer-hours to operate your own. Treat internal service-to-service encryption as a feature your platform provides, not a project you staff.

Enterprise. With hundreds of services across many teams and languages, a mesh earns its complexity, and the real work is governance: one written division of labour so gateway and mesh never double-handle a concern, uniform mTLS and telemetry, and a platform team that owns proxy upgrades and certificate rotation. Auditors want a single provable enforcement point for access and encryption, so standardise the interface and make policy something teams inherit rather than reimplement. Manage the sidecar tax as a real budget line, and keep sidecarless options open so you are not locked out of cheaper approaches later.

Government. Procurement rules, transparency, and public accountability shape the architecture. Treat the gateway and mesh as the zero-trust backbone regulators expect: every citizen and partner authenticated at the front door, every internal call authenticated and encrypted by workload identity, and an audit trail at each enforcement point proving where access was checked and where traffic was encrypted. Favour open standards and portable configuration over proprietary lock-in that procurement rules may forbid, and publish, where appropriate, how the platform protects citizen data as it crosses each boundary.

Examples

Startup. A twelve-person startup runs eight services behind a single API gateway. The gateway handles all the north-south work: it authenticates users with a token check, enforces per-plan rate limits so free-tier users cannot swamp the system, and composes a couple of chatty endpoints into single mobile-friendly calls. For east-west traffic they deliberately skip a service mesh, because eight services in two languages do not justify a control plane. Instead they get mTLS and retries from a shared client library and their platform’s built-in certificate management, and they collect traces with a lightweight agent. When they later add a dedicated mobile client with tighter payload needs, they introduce a mobile backend for frontend beside the existing web gateway. They revisit the mesh question every year and keep deciding, correctly, that they have not yet crossed the threshold where it would pay.

Enterprise. A multinational bank runs several hundred services across many teams and languages, and here a service mesh earns its complexity. Every service gets a workload identity and mTLS by default, so all internal traffic is authenticated and encrypted without any team writing crypto code, which satisfies both the security organisation and the auditors who want a single provable enforcement point. The mesh applies uniform retries, timeouts, and circuit breaking by policy, and it shifts traffic gradually for canary releases so a bad deploy touches one percent of users before it touches all of them. A tier of API gateways fronts the external and partner traffic, owning authentication, quotas, and versioning, with a firm written rule that internal retries live only in the mesh and external rate limits live only in the gateways, so the two layers never double-handle a request. Consistent telemetry from every proxy feeds a central observability platform that lets one on-call engineer trace a request across dozens of service hops.

Government. A national tax agency modernises a citizen-facing filing platform and treats the gateway and mesh as the backbone of a zero-trust architecture that regulators require. An API gateway is the controlled front door: every citizen and every partner is authenticated and authorised there, external rate limits protect the system during the filing-deadline surge, and old client formats are transformed at the edge so legacy integrations keep working. Behind it, a service mesh gives every internal service a cryptographic identity and enforces mTLS on every call, so no service is trusted merely for being inside the network, and access policy explicitly lists which services may call which. Because the platform spans multiple data centres for resilience, the mesh extends the same identity and encryption guarantees across clusters, giving consistent zero-trust networking nationwide. Every enforcement point emits an audit trail, so the agency can prove to oversight bodies exactly where access was checked and where traffic was encrypted.

Business case: motivations, ROI, and TCO

The return on a gateway is easy to see and usually large. Instead of every service reimplementing authentication, rate limiting, and request logging, you build those once at the edge and every service inherits them. A security fix or a new rate-limit policy ships in one deploy rather than a hundred, which shortens the time to close a vulnerability and lowers the cost of every audit, because there is a single place to inspect. The composition and versioning features cut client round trips and let you evolve APIs without breaking callers, which reduces both latency costs and the coordination tax between teams. For almost any system with external clients, a gateway pays for itself quickly.

The service mesh has a subtler business case, because its total cost of ownership is real and ongoing. You pay for the proxies’ compute and latency, for the engineers who operate the control plane, and for the learning curve of debugging a new layer. That cost is justified when the alternative (per-service, per-language implementations of mTLS, retries, and telemetry) would be larger and, worse, inconsistent in ways that create security gaps and outages. At high scale the mesh converts a hundred fragile hand-rolled solutions into one uniform, provable one, and the ROI shows up as fewer security incidents, faster safe deployments through traffic shifting, and dramatically better observability. Below that scale, the honest calculation often favours libraries and a gateway, and the disciplined move is to wait. To make the case to leadership, connect the gateway to the metrics they already track (time to remediate vulnerabilities, audit cost, API coordination overhead) and connect the mesh to security incident containment, deployment safety, and the cost of the polyglot duplication it replaces.

Anti-patterns and pitfalls

  • Mesh before you need it: adopting a full service mesh at a handful of services, buying a control plane to solve a problem you do not yet have.
  • Double retries: the gateway and the mesh both retry, so one client call multiplies into a backend retry storm that amplifies an outage.
  • Timeout inversion: an inner timeout longer than the outer one, so the caller gives up while the callee keeps working on a response no one will read.
  • Gateway as a monolith: stuffing business logic into the gateway until it becomes a shared bottleneck that every team must coordinate to change.
  • Single point of failure: running one gateway instance with no redundancy, so the front door failing takes down every service behind it.
  • Trusting the network: treating anything inside the perimeter as safe, so one compromised service can move laterally and reach everything.
  • Overlapping ownership: no written division of labour, so the same concern is handled in both layers by accident and no one knows which is authoritative.
  • BFF sprawl: spinning up a backend for frontend per client when the clients want the same thing, multiplying gateways and duplicating logic for no gain.
  • Ignoring the sidecar tax: deploying thousands of sidecars without measuring the compute and latency overhead, then wondering where the budget went.

Maturity model

  • Level 1, Initiate: Services talk to each other directly with no shared layer. Authentication, retries, and timeouts are hand-coded per service and inconsistent. Internal traffic is often plaintext and trusted for being on the network, and there is no single place to enforce policy or observe traffic. Decisions are reactive, made service by service as problems surface.
  • Level 2, Develop: An API gateway fronts external traffic and centralises authentication, rate limiting, and routing, but east-west concerns are handled by shared libraries with uneven adoption. Some services have mTLS; many do not. Teams recognise the north-south versus east-west distinction, yet ownership is informal and practices differ from one team to the next.
  • Level 3, Standardise: North-south and east-west concerns are cleanly separated with a written division of labour applied across the organisation. The gateway owns external identity, quotas, and composition; a mesh or a consistent library layer owns internal mTLS, retries, and telemetry. Overlaps are resolved so no concern is double-handled, internal traffic is encrypted and identity-checked by default, and the standard is documented and enforced rather than left to each team’s discretion.
  • Level 4, Manage: The traffic layer is measured and controlled against baselines. You track the sidecar tax in compute and tail latency per hop, gateway and mesh availability against error budgets, retry-amplification and timeout-inversion incidents, mTLS coverage as a percentage of internal calls, and the time to push a policy change across the fleet. Metrics gate decisions: a proxy upgrade, a new retry budget, or a canary rule is judged on data against a baseline rather than on intuition, and any drift from the standard triggers a fix.
  • Level 5, Orchestrate: Gateway and mesh are policy enforcement points for a mature zero-trust architecture, with identity checked on every hop across clusters and regions. Traffic shifting drives safe progressive delivery, observability is uniform and rich, and the organisation continuously evaluates and adopts sidecarless, proxyless, and eBPF approaches where they pay. The traffic layer is integrated with security, delivery, and capacity planning, and it adapts as scale, languages, and the risk picture shift.

Ideas for discussion

  1. If your single API gateway went down right now, how many services would become unreachable, and what is your plan for making the front door redundant?
  2. Which concern in your system is currently handled in both the gateway and a service (or a library), and how would you prove it is not being double-handled?
  3. At what number of services and languages would your team agree that a mesh has finally earned its complexity, and how far are you from that line?
  4. If an attacker compromised one internal service tomorrow, which other services could it reach, and what identity check would stop it?
  5. Are the sidecarless or eBPF-based mesh approaches mature enough for your platform yet, and what would you measure to decide?
  6. Does each of your divergent clients genuinely need its own backend for frontend, or are you about to duplicate logic that could stay shared?

Key takeaways

  • Separate north-south traffic (handled by an API gateway) from east-west traffic (handled by a service mesh); each owns its direction, and confusing them causes double-handling.
  • A gateway centralises routing, authentication, authorisation, rate limiting, quotas, request transformation, composition, and versioning, so services stay thin and policy lives in one place.
  • A service mesh gives you mTLS, traffic shifting, platform-level retries and circuit breaking, and uniform observability without changing service code, classically through sidecar proxies.
  • Adopt a mesh only when the count of services and languages makes per-service wiring the more expensive path; below that, libraries plus a gateway win, and the sidecar tax is real enough to watch.
  • Treat gateway and mesh as the policy enforcement points of a zero-trust architecture, assign each cross-cutting concern to exactly one layer, and check identity on every hop across clusters.

References and further reading

  • Sam Newman, Building Microservices: Designing Fine-Grained Systems
  • Chris Richardson, Microservices Patterns: With Examples in Java
  • Lee Calcote and Zack Butcher, Istio: Up and Running
  • Ken Owens, Alois Reitbauer, and others; the CNCF Cloud Native Landscape and service mesh documentation
  • Evan Gilman and Doug Barth, Zero Trust Networks: Building Secure Systems in Untrusted Networks
  • Scott Rose, Oliver Borchert, Stu Mitchell, and Sean Connelly, Zero Trust Architecture, NIST Special Publication 800-207
  • Susan Fowler, Production-Ready Microservices