Engineering at WebMD Ignite · 1/1
11 seconds to 500 ms: making a healthcare platform faster by adding endpoints
WebMD Ignite case study: 30 overloaded endpoints became 46 focused ones, and screens, indexes and caching followed. Responses went from 3–11 s to 150–500 ms.
On this page
For almost two years I worked on an internal platform for WebMD Ignite: software that healthcare staff used to manage care information, assets and asset groups. Critical software, with a lot of personal data behind every screen. I was one of about 25 people on the team, working with QA, product managers, a product owner, a tech lead and my team lead.
This is the story of the part I’m proudest of: making the slowest parts of the app fast, without a rewrite.
The system
A microservices architecture with a .NET gateway in front, mostly REST with some older parts still on GraphQL, a Vue single-page app for the main product, and some older sites still running on a PHP CMS with React islands. PostgreSQL held the core data, Elasticsearch and Solr powered search and some of the services’ reads, and RabbitMQ carried events between services. An ETL process turned an old database, which stored records as XML, into properly structured data in Postgres, which was then indexed for search. Background jobs ran on Hangfire, files lived in AWS S3 buckets, and Cloudflare sat in front. The front end was Vue with Tailwind, tested with Vitest and Cypress.
Changes went through dev, QA, staging, a pre-production environment and production, with Grafana, Kibana, Datadog and Sentry to watch what happened once they landed.
The problem
Most screens waited around 3 seconds for their data, and the worst endpoints took up to 11 seconds. For people who use these screens all day, often with someone waiting on the answer, that’s not a detail.
When I dug in, the cause wasn’t one slow query. It was the shape of the API. About 30 endpoints had grown to do far too much: each returned everything a screen might ever need, joined across several tables, whether the screen used it or not. Every request paid for the worst case.
What that looked like
Picture a screen that lists the asset groups of a department, showing each group’s name and how many assets it has. The endpoint behind it looked like this:
GET /asset-groups?department=42
[
{
"id": 7,
"name": "Imaging equipment",
"assets": [ { "...": "every asset, with every field" } ],
"careRecords": [ { "...": "every care record linked to those assets" } ],
"auditTrail": [ { "...": "the full history of changes" } ]
}
]
The screen used two fields per group. It received every asset, every linked record and the full audit history, for every group in the department, in one response. After the change:
GET /departments/42/asset-groups?page=1&size=25 → id, name, assetCount
GET /asset-groups/7/assets?page=1&size=25 → only when a group is opened
GET /asset-groups/7/audit?page=1&size=25 → only when the history tab is opened
The list screen now makes one small call. The details load only when someone asks for them.
The database side followed the same logic. A typical case:
SELECT id, name, asset_count
FROM asset_groups
WHERE department_id = 42
ORDER BY updated_at DESC
LIMIT 25;
- An index on a
statuscolumn with three distinct values was almost never picked by the planner, and it slowed down every write, so it went. - A composite index on
(department_id, updated_at)matches both the filter and the sort, so the planner reads 25 rows in order instead of scanning and sorting the table.
What I changed
1. Endpoints that each do one job
I split those 30 endpoints into 46 focused ones. Each returns what one view actually needs, in a response shaped for it. More endpoints sounds like more work for the client, but each call is small and fast, and screens can load their parts in parallel instead of waiting for one giant payload.
The payloads changed too. The old responses were big DTOs built to carry data for a particular front end: whatever the screen might want, pre-assembled, whether it used it or not. The new ones are leaner and more generic: each resource comes back with the fields it actually has, in a shape any consumer can use and combine. The front end asks for what it needs instead of the API guessing.
And where an endpoint returned a list, I added pagination. That cut response time and payload size together: a screen showing the first 25 items no longer downloads, serialises and parses thousands.
I documented the new endpoints and the patterns behind them, so the next person adding a screen would reach for a focused endpoint instead of growing an old one.
2. Screens that ask for less
Splitting the endpoints exposed a design question: some screens wanted all the information at once, so no endpoint could be fast for them. So we redesigned them. The main screen became a shallow overview that loads instantly, and the detail moved into sub-screens that fetch exactly what they show, only when someone opens them.
That only worked because the design, the product owner and the engineers agreed on it. An API can’t be faster than the screen it serves: if the design asks for everything up front, the backend has to deliver everything up front. Design and architecture have to be aligned.
3. Indexes that earn their place
With the endpoints reshaped, the remaining time was in the database. I went through the indexes with the query plans in hand:
- Dropped indexes the planner rarely used. They cost time on every write and bought almost nothing on reads.
- Added indexes where the queries needed them, checking column cardinality first, so an index would actually narrow the search instead of being ignored.
- Normalised or kept redundancy on purpose, case by case. Some data was better normalised; some read-heavy paths were better served by a little deliberate redundancy.
On the .NET side the code was async all the way down, so a request waiting on the database or another service never held a thread. Cancellation tokens were passed through to the database and downstream calls, so when the front end aborted a request, the server stopped working on it too instead of finishing a query nobody would read. Responses were gzip-compressed, which matters when a list page ships a lot of JSON.
4. Caching on both ends
Focused endpoints are also easier to cache, because each one has a clear owner and a clear shape. We cached on the server side, with .NET’s caching, for data that was read far more often than it changed, and on the client with TanStack Query, so moving between screens reused fresh data instead of asking for it again.
5. A lighter front end
On the Vue side I refactored the UI around proper patterns:
- No request waterfalls. Calls that didn’t depend on each other were fired together (
Promise.all) instead of one after another, so a screen waited for its slowest call, not the sum of all of them. - Skeletons instead of spinners. Each part of a screen showed its shape straight away and filled in as its data arrived, so the page felt ready before it was.
- Debounced inputs, so typing in a search box didn’t fire a request per keystroke, and cancelled requests (
AbortController), so leaving a screen or typing again dropped the old call instead of letting it finish for nothing. - Stale-while-revalidate with TanStack Query: cached data showed instantly and refreshed in the background.
- Prefetching and lazy loading where they paid off: a sub-screen’s data could be on its way before the click, and heavier screens loaded their code only when opened.
- Virtualised long lists, for infinite scroll and other long views: only the rows on screen were rendered, however long the list.
- Logic in composables, components split into smart and dumb. Fetching, state and business rules moved into composables, which Vitest can test on their own without mounting anything. A smart component per screen orchestrates them; the dumb components below it only receive props and render. Easier to test, easier to reuse, and a smaller bundle, because the same pieces served many screens.
Shipping it without breaking anything
A live, critical app depended on the old endpoints, so nothing was switched off overnight. The new endpoints shipped as a new version next to the old ones, and the old ones were deprecated gradually: consumers moved over a piece at a time, and after each step we watched Datadog, Sentry and the rest of our observability for errors before moving the next. Only when nothing depended on an old endpoint any more did it go away.
The result
| Before | After | |
|---|---|---|
| Typical response | ~3 s | 150–250 ms |
| Worst endpoints | up to 11 s | ~500 ms |
| Endpoints | 30, doing too much | 46, one job each |
| Payloads | front-end DTOs, full lists | lean, generic resources, paginated |
| Speed-up | roughly 12 to 22× |
Same data, same infrastructure, no rewrite: the gain came from asking each request to do only the work it needed.
The trade-offs
Everything in architecture is a trade-off, and this change was no exception.
- What it protected: performance above all, because people used these screens all day, often with someone waiting on the answer. Maintainability came with it: an endpoint with one job is easier to read, test and change than one that serves every screen.
- What it cost: more endpoints to document and version, 46 instead of 30. Screens that used to make one call now make a few. Server-side caching trades a little freshness for speed, so it only went on data that was read far more often than it changed.
- What had to stay safe: availability. Two versions of the API lived side by side for a while, and that only works if you can see what’s happening, which is why every step was checked in Datadog and Sentry before the next.
Spot it in your own app
- Open DevTools, reload a slow screen, and sort the Network tab by size and time. If an endpoint returns much more than the screen shows, or a list the screen then pages through itself, it’s doing too much.
- Run the slowest query behind it with
EXPLAIN ANALYZE. A sequential scan on a column you filter and sort by is usually a missing composite index. - Count how many screens call the same endpoint. If they all use different parts of the response, it’s several endpoints in one.
Takeaways
- Shape the API around what each screen needs, not around everything it might need.
- Design and architecture have to be aligned. A screen that asks for everything forces an endpoint that returns everything.
- Every index has a cost. Keep the ones the planner uses and the queries need; drop the rest.
- Paginate lists by default. It saves time and bytes on both ends.
- Measure before and after. The numbers are what turn a refactor into a decision the whole team can agree on.
The hard part
The legacy side. Parts of the system still ran on the PHP CMS, and it became the bottleneck: the older code and runtime set limits we couldn’t optimise past from the outside. Some of the biggest wins were about working around it without breaking the sites that depended on it.
How we worked
I pushed for good practices while keeping features shipping on time, which is mostly a balancing act: knowing when a refactor pays for itself this sprint and when it should wait. Knowledge sharing happened in code reviews and in one to three pair-programming sessions a day.
What I’d do differently
More deliberate knowledge sharing about the architecture itself. Reviews and pairing spread a lot, but they spread it one person at a time. Regular short sessions on how the services fit together, and on the design patterns we’d agreed on, would have made everyone more comfortable across the whole codebase, not just the parts they touched.
Thanks
A big hug to the Rebel Alliance team, and to my team lead and good friend Cory Cookson.