Software Design · 2/3

From modules to agents: how architecture evolved, and how to measure modularity

A short history of software architecture, from Dijkstra's layers to cloud, containers and LLM agents, then modularity vs granularity, cohesion types and the LCOM metric.

On this page

The first post was about the job. This one steps back to see how the field got here, then zooms in on the idea that has survived every era: modularity, and how to tell a good module from a bad one.

A short history

Thinking in layers and modules

The earliest architecture thinking was about decomposition: splitting a system into modules that hide their internals behind interfaces, and stacking them in layers, where each layer uses only the one below. Both ideas are still everywhere, from the OSI network model to the controller–service–repository layering of most web backends.

They’re older than the word “architecture” itself:

WhenWhatWhy it mattered
1968Dijkstra’s THE operating systemOne of the first systems described as a stack of layers, each built only on the one below.
1968NATO Software Engineering ConferenceCoined “software engineering” in answer to the software crisis: projects late, over budget and unreliable.
1972Parnas, information hidingSplit a system by the decisions likely to change, and hide each one inside a module behind its interface. The public/private boundary we take for granted starts here.
1992Perry and Wolf“Software architecture” becomes a named subject: elements, form, and rationale.
1995Styles and viewsThe first catalogues of styles (client–server, layered, pipe-and-filter, the Unix pipe being the famous one), and Kruchten’s 4+1 view model: logical, process, development and physical views, tied together by scenarios.
1996Shaw and GarlanSoftware Architecture: Perspectives on an Emerging Discipline. Architecture as a discipline of its own.
1998Software Architecture in Practice, 1st ed.Architecture as styles and components. Decisions and quality attributes come to the centre in later editions.
~2000Patterns, SOA, quality attributesThe POSA pattern books; service-oriented architecture and web services (SOAP and XML, later REST, later gRPC); quality attributes and documentation (Clements and colleagues) become part of how architecture is done.
late 2000sEnterprise architectureArchitecture at the scale of a whole organisation, with its own frameworks and certifications.

Architecture description languages

In the 1990s, research tried to make architecture formal. Architecture description languages (ADLs) like Wright, Darwin, Rapide, Acme and later AADL described components, connectors and configurations precisely enough to analyse them: check for deadlocks, verify that a system matched its style.

They did describe architecture well. They just never became part of everyday industrial practice: too heavyweight, too far from the code, and out of date as soon as the code moved on. What survived were the ideas, especially components and connectors as the vocabulary for runtime structure, which Documenting Software Architectures builds on.

Today the closest thing in daily use is diagrams as code, like Mermaid in a README: versioned next to the code and easy to update, but it draws parts of an architecture, it doesn’t specify one. There’s still no widely used formal language for the whole thing.

Enterprise architecture vs software architecture

The two are easy to mix up:

Enterprise architectureSoftware architecture
ScopeThe whole organisationOne system, or a few related ones
ConcernAligning business processes, information, applications and technologyThe structure and qualities of the software
Typical questionsWhich systems do we need, how do they share data, what do we buy or build?How is this system split, how do its parts communicate, how does it scale?
FrameworksTOGAF, ZachmanViews and styles, quality attributes, ADRs

An enterprise architect decides that the company needs a customer data platform; a software architect decides how that platform is built.

The cloud changes the questions

Then infrastructure became software, and the characteristics that were expensive before became cheap to get:

WhenWhatWhat it changed
2006Amazon EC2Servers on demand, paid by the hour. Scalability and elasticity stop needing hardware purchases.
~2009DevOpsDevelopment and operations as one team and one pipeline: build, test, deploy and run continuously.
2011The Twelve-Factor AppA checklist for cloud-ready services: config in the environment, disposable processes, logs as streams. Operations concerns written down as design rules.
2011ADRs (Michael Nygard)Lightweight decision records kept in the repo, next to the code they explain.
2013DockerContainers: package a service with everything it needs and run it the same way everywhere. The end of “works on my machine”.
2014KubernetesOrchestrating containers at scale: scheduling, restarts, rolling deploys, scaling. Now mostly consumed as a managed service on AWS, Azure or GCP.
2014AWS LambdaServerless: run functions, not servers, triggered by events (a request, a file upload, a row inserted), and pay per call.
~2014MicroservicesMany small, independently deployable services, each owning its domain and its data, loosely coupled. Only practical because of everything above.
2016Site Reliability Engineering (Google)Operations as an engineering discipline with its own patterns and metrics: SLOs, error budgets, toil. DevOps with numbers.
2017Istio (service mesh)Moves service-to-service concerns (security, monitoring, retries, timeouts, circuit breaking) out of the code and into the platform.
2017Evolutionary architectureFitness functions: automated checks that keep a system conforming to its quality attributes as it changes.
2020RAGRetrieval-augmented generation: a language model answers using documents fetched for the question, not only what it memorised.
2022ReActA loop where a model reasons about the task, acts (searches, calls a tool), observes the result, and reasons again. The “thinking” you see in today’s assistants is this loop made visible.
2022LangChainFrameworks that abstract over models and wire them to tools and data.
2023–AgentsAn agent is an LLM plus tools plus memory: less a new kind of program than a configuration. Then agents that spawn other agents, and frameworks to orchestrate them.
late 2024MCP (Model Context Protocol)A standard way for agents to reach tools and data. Adopted across competing vendors in months, which is remarkably fast for a standard.
2025Agent-to-agent protocols (A2A)A standard for agents talking to other agents, not only to tools.
nowHarnesses and human-in-the-loopPatterns for bounding what agents can do, with people designed in as components of the architecture: approving, correcting, taking over.

Each step moved a hard problem into the platform, and each one created new architectural decisions in its place. Microservices make scaling individual parts easy and data consistency hard. Agents make open-ended tasks possible and predictable behaviour hard. The first law holds.

What didn’t change

The lecture closed on the other side of the timeline: what stayed the same through all of it.

  • Architecture is the set of structures that let people reason about a system. Layers in 1968, services in 2014, agents and tools now.
  • There’s no right architecture, only trade-offs that fit a context better or worse.
  • Quality attributes drive the decisions. The cloud made some cheap; it didn’t make choosing unnecessary.
  • Documentation still matters. It moved from formal languages to ADRs and diagrams as code, but the why still has to be written down.
  • Context decides everything. The same choice is right in one system and wrong in the next.

Modularity

Through all of those eras, one idea stayed central. Richards and Ford describe modularity as the fundamental organising principle of architecture, and for a reason: software systems tend toward entropy. Left alone, structure erodes, responsibilities blur, and everything ends up depending on everything. Keeping a system modular takes continuous effort, and to spend that effort well you need a way to judge a module.

Two things make this harder than it sounds. No client will ever ask for it: nobody writes “the system must be modular inside” in a requirements document, so it’s on the architect to care. And while countless books say modularity is good, far fewer say how to split a system.

A sandwich shop

The course works through an example from Richards and Ford: an online sandwich shop. Customers order in the shop, online or at a kiosk, and the system has to:

  • show the menu and prices, with discounts;
  • let customers customise a sandwich (extra toppings, a different bread);
  • place and validate the order;
  • take payment and apply promotions;
  • route the order to one of several kitchens, since a big shop might make thousands an hour;
  • notify the customer when the sandwich starts and when it’s ready;
  • for online orders, assign a courier.

The first design that comes to mind, and the way plenty of systems were actually built, is one giant OrderManager that does everything: reads the menu, calculates prices, charges the card, updates the kitchen, sends notifications. It works, and it’s a problem. All the code is in one place, so a change to payments means editing the same class as a change to delivery, and every change risks everything else.

Splitting it into a module per responsibility (menu, orders, payments, kitchen, notifications, delivery) is the intuitive fix. But how far to split is a real question. A kitchen module with separate submodules for slicing the bread and toasting it? Probably not; there’s no business logic to separate. A payments module with a submodule per payment method (card, MB WAY, PayPal)? That one might be worth it, because new methods will come and each has its own rules.

Modularity vs granularity

That’s the key idea: the more modular a system, the smaller its pieces, and past a point, small pieces cost more than they save. Too fine-grained, and modules spend their time calling each other, a change needs several of them to move together, and a distributed system can end up as a distributed monolith: all the coupling of a monolith, plus the network in between. The book’s rule of thumb: embrace modularity, but watch the granularity.

Modules will still need to talk: if a payment fails, the order has to be cancelled and the stock restored. That communication is coupling, and the goal isn’t zero, it’s as little as the design allows. The fewer points of contact between modules, the easier each one is to change and to test.

One more distinction: a module is logical, not physical. Two modules can live in the same package, the same repository, even the same deployable, and still be separate modules, as long as their boundaries are respected. A modular monolith is exactly that. Packages, JARs and .NET assemblies are something else: units of deployment, decided by how the system is published (which machines, which containers). A module can map to one package, or many modules can ship in one; the two decisions are related but separate.

Cohesion

The main measure of a module is cohesion: how much its parts belong together.

From best to worst

CohesionThe parts are grouped because…Example
Functional (best)every part contributes to one well-defined task, and the module has everything that task needsa PriceCalculator that computes a price, and nothing else
Sequentialthe output of one part is the input of the nextUnix commands, designed so one’s output can be piped into the next; or parse → validate → save in one import module
Communicationalthe parts work on the same data, or contribute to the same outputall the code that touches one database table in one module: common in older systems, where data access shaped the structure more than the business logic did
Proceduralthe parts must run in a particular order, though they share little data“check permissions, then log, then send the email”
Temporalthe parts run at the same time, not because they’re relatedthe start-up code of an app or an OS: an init() that opens the database, loads config and warms the cache
Logicalthe parts do related kinds of thing, but not the same thinga StringUtils class, or a utils.ts with unrelated helpers
Coincidental (worst)nothing; they just ended up in the same placea Misc module

The higher the cohesion, the easier a module is to understand, test, reuse and change, because a change to one task touches one module. The order isn’t a ban on the lower ones: in some architectures, grouping by data or by sequence is the right call. It’s a preference.

Two pairs are easy to mix up. Sequential and procedural both run steps in order, but in sequential cohesion each step’s output is the next one’s input, like a Unix pipe; in procedural cohesion the order matters and nothing flows between the steps. And logical and coincidental both look like a drawer of unrelated things, but logical cohesion has a rule, however weak (“string helpers”, or even “every method starting with A”), while coincidental has none at all.

The sandwich shop’s giant OrderManager is a good test. Everything in it is “about orders”, which sounds cohesive, but the only thing holding payments, kitchen routing and delivery together is that loose theme. That’s logical cohesion at best.

Two honest notes from practice:

  • Logical cohesion isn’t always a crime. A file of small, pure utilities for one domain, like colour helpers for a theme, is logically cohesive and perfectly fine. It becomes a problem when a generic utils grows into a drawer where anything goes, which is coincidental cohesion wearing a better name.
  • Functional cohesion is what good composables and components aim for. When I built a company-wide form-input library, each composable owned exactly one concern (keyboard navigation, focus and scrolling, accessibility), and the components only composed them. That’s functional cohesion by design.

Measuring it: LCOM

Cohesion can be measured, at least for classes. The classic metric is LCOM, lack of cohesion in methods (Chidamber and Kemerer, 1994). The idea: methods of a cohesive class use the same fields; methods that share nothing probably belong to different classes.

The original version counts pairs of methods:

  • PP = the number of method pairs that share no fields;
  • QQ = the number of method pairs that share at least one field;
LCOM=max⁡(P−Q, 0)\mathrm{LCOM} = \max(P - Q,\ 0)

Zero means cohesive; the higher it gets, the more the class looks like several classes glued together. Take this one:

class Customer {
  private name = '';
  private email = '';
  private orders: Order[] = [];
  private cart: Item[] = [];

  rename(name: string) { this.name = name; }                 // uses name
  changeEmail(email: string) { this.email = email; }         // uses email
  addOrder(order: Order) { this.orders.push(order); }        // uses orders
  totalSpent() { return sum(this.orders.map((o) => o.total)); } // uses orders
  addToCart(item: Item) { this.cart.push(item); }            // uses cart
}

Five methods make (52)=10\binom{5}{2} = 10 pairs. Only one pair shares a field (addOrder and totalSpent, through orders), so Q=1Q = 1, P=9P = 9, and LCOM=8\mathrm{LCOM} = 8. The number is telling us what reading the class suggests: it’s three classes, a profile, an order history and a cart, sharing a name. Split them, and each one scores 0.

LCOM has known weaknesses. It ignores methods that call each other, it’s thrown off by getters and setters, and several later variants (LCOM2 to LCOM4, and others) fix different flaws. It’s a smell detector, not a verdict. It’s also the kind of formula that code-analysis and refactoring tools compute for you: you’ll rarely calculate it by hand, but it’s worth knowing that “this class should be split” can be backed by a number and not only by taste. Alongside it, Richards and Ford cover coupling metrics: how many modules depend on a module (afferent coupling) and how many it depends on (efferent coupling), and derived measures like abstractness and instability.

What I’m taking from this

  • Every era moved a hard problem into the platform, and created new decisions in its place. The tools change; trade-offs stay.
  • Formal description languages failed in practice, but their vocabulary won. Components and connectors are still how we talk about runtime structure.
  • The tools changed; the fundamentals didn’t. Structures to reason with, trade-offs, quality attributes, documentation, context.
  • Modularity is maintenance against entropy. It doesn’t stay by itself, and nobody will ask you for it.
  • Split, but watch the granularity. Too many tiny modules is its own kind of mess, and a distributed monolith is the worst of both.
  • Judge a module by why its parts are together. Aim for functional cohesion; be suspicious of “utils”.
  • Metrics like LCOM point, they don’t decide. A high score is a reason to look, not a reason to split.