Software Architecture & Systems Design

Can a 20-Person Team Trust This Architecture Framework?

A framework can make a microservices demo look convincing while leaving a small engineering team with an expensive operating model. For a 15–30 person organisation, I would favour vendors that solve one painful boundary and can be removed without rewriting the application. I would not buy a comprehensive “microservices platform” on architectural promise alone, because its maintenance burden arrives long before its advertised scale does.

The vendor should earn one boundary, not inherit the whole system

Begin the evaluation with the boundary that currently causes trouble: perhaps synchronous calls between two services fail together, or publishing an event safely requires code copied across repositories. Read Scalable Software Architecture Patterns for Modern Systems as a set of design hypotheses rather than a purchasing specification, because a pattern is useful only where your failure modes justify its moving parts. Ask each vendor to show how its product improves that particular boundary without requiring changes to unrelated services.

Make the candidate implement the same small workflow twice: once with its recommended approach and once with tools you already operate. A useful trial might accept an HTTP request, write an order to PostgreSQL, publish an event, and let another service consume it idempotently. Require the team to demonstrate what happens if the database commits but publication fails, if the consumer restarts midway through processing, and if the message arrives twice. A polished successful path is weak evidence, because most architectural cost sits in recovery and diagnosis.

Compare MassTransit with RabbitMQ against Dapr pub/sub with RabbitMQ rather than comparing either with a slide labelled “event-driven architecture.” MassTransit can win when your team wants messaging behavior visible in .NET code and is comfortable owning broker configuration; it costs developer attention for retries, consumer conventions, and version upgrades. Dapr can win when several language stacks need a common sidecar interface; it costs another deployed process, component configuration, and a larger failure surface. Neither option removes the need to decide who owns event schemas or duplicate handling.

Write down which responsibilities remain yours after purchase. For the trial workflow, these include the PostgreSQL transaction boundary, message identity, retry policy, poison-message handling, and replay procedure. If a vendor says its framework “guarantees delivery,” ask whether that means broker persistence, at-least-once delivery, or atomicity with your database, because those are different guarantees. Reject any answer that depends on a proprietary dashboard without an exportable event history; an incident should remain intelligible if the subscription ends.

A reference architecture is evidence only when your team can operate it

Have two engineers who did not build the trial deploy it from an empty environment using the candidate’s documentation. Time their work, but also record each undocumented decision and every credential they must request. That exercise tests the product your team will actually consume, whereas a vendor-assisted installation tests the vendor’s solutions engineer. Give the reviewers access to the same infrastructure and permissions you would allow in production; special administrative access can conceal an otherwise difficult operating model.

Demand an observable failure, not a dashboard tour. Kill a consumer during processing, interrupt RabbitMQ connectivity, then ask the reviewers to find the affected request through W3C Trace Context, OpenTelemetry traces, and broker metadata. OTLP export to an OpenTelemetry Collector is preferable to an agent that only feeds the vendor’s UI when your team already uses Grafana or another tracing backend, because it keeps incident investigation portable. Check whether trace identifiers survive asynchronous messages; an HTTP-only trace can make the hardest part of the workflow disappear.

Review Designing Scalable NET Microservices Architecture at Scale against this operating test rather than treating microservice decomposition as a reason to adopt more infrastructure. A framework that requires a separate control plane, service mesh, and custom deployment operator for one order workflow has to justify each dependency, because every additional component needs an owner during an outage.

Ask for the exact supported versions of .NET, PostgreSQL, RabbitMQ, and Kubernetes, plus the vendor’s policy when one of them reaches end of support. Microsoft publishes a three-year support term for .NET 10 as an LTS release; that published window is useful context, but it does not establish that a candidate library will support your next runtime upgrade promptly. Put the vendor’s promised upgrade interval in writing, then verify its last two releases against public changelogs and issues. A release cadence is credible when previous compatibility work is visible, not when a sales document calls the product “future-proof.”

The trial should price failure and exit, not merely throughput

Run the candidate and your current implementation under the same workload, with identical PostgreSQL data, broker durability settings, and resource limits. The following k6 script is a runnable starting point for an HTTP endpoint on a locally running service; change BASE_URL to test each implementation. Its thresholds are prompts for discussion, not claims about your production requirements.

import http from 'k6/http';
export const options = {
  vus: 20,
  duration: '60s',
  thresholds: {
    http_req_failed: ['rate<0.01'],
    http_req_duration: ['p(95)<250'],
  },
};
export default function () {
  http.get(__ENV.BASE_URL || 'http://127.0.0.1:5000/health');
}

The 20 virtual users for 60 seconds are an initial trial setting to tune until the request mix resembles your traffic; a health endpoint alone cannot validate event processing. Treat 250 ms at the 95th percentile as a provisional limit to negotiate with the product owner, since an endpoint’s purpose determines whether that latency matters. If your own trial measures 180 ms for the current path and 310 ms for the candidate under equivalent conditions, those observed values warrant investigation rather than an automatic rejection: the candidate may be doing durable work the current path skips. Record completed business operations, queue depth, and time to recover alongside HTTP latency so the comparison does not reward lost work.

Price the deployment required to keep those results. Include broker storage, sidecar memory, telemetry volume, paid support, and engineer time spent on configuration. If the vendor charges per message, model a retry storm as well as normal traffic, because failures can increase billable attempts precisely when the team has the least time to control them. Ask whether a test environment needs the same licensed components as production; otherwise, the supposedly low entry price may discourage realistic rehearsals.

Make exit a trial task. Replace the candidate’s client with a small adapter, replay a saved batch of events, and export the data needed to continue elsewhere. Prefer CloudEvents 1.0 or a documented JSON Schema for exchanged events where they fit, because independently readable contracts make replacement less dependent on vendor code. Count the source files and deployment manifests changed during removal, then ask whether that change is feasible within one planned iteration. If not, price the lock-in explicitly instead of calling the integration “lightweight.”

Contract terms matter only when they match the technical escape route

Procurement can protect engineering only when its promises correspond to testable behavior. Ask who owns a production incident involving the framework and RabbitMQ: the vendor, your team, or a third-party broker provider. Require a support escalation path that names the logs and diagnostics the vendor needs, because a response-time promise is less useful if the first response asks for telemetry you cannot collect. Verify that support covers the deployment model you tested rather than only the vendor’s hosted service.

Inspect licensing at the dependency level. Check the library’s license, any commercial extensions used in the reference implementation, and whether the NuGet package has transitive dependencies with different terms. Enable NuGet lock files through RestorePackagesWithLockFile and retain packages.lock.json for the trial, because an evaluation you cannot reproduce makes both security review and later cost estimates less reliable. Look for a documented vulnerability disclosure process and a maintained issue tracker; a security questionnaire cannot substitute for evidence that fixes reach supported versions.

Put the exit exercise into the decision record before signing. Specify that the team must retain its event schemas, operational telemetry, and a working migration procedure if the agreement ends. I would not accept an exclusive event format merely to save a few days of integration work, because it turns an ordinary renewal negotiation into a data migration project. A vendor willing to help demonstrate an export path is offering stronger evidence than one willing only to promise portability.

The first decision should be a small, reversible experiment

Choose one boundary that has produced a real incident or repeated engineering work, then write its failure scenario and acceptance criteria before contacting vendors. Assign one engineer to build the current-tool baseline and another to test a candidate against the same workflow. Book the exit rehearsal on the calendar now. If the candidate cannot survive that narrow test, its broader architecture is not yet your team’s problem.