Event-driven architectures are powerful, but they’re also complicated. You’ve got event schemas that drift. Producers and consumers that get out of sync. Dead letter queues nobody monitors. Idempotency handled inconsistently. Ordering guarantees that work in theory but fail in practice. Claude Code can automatically audit all of this, generating comprehensive architecture reviews that spot real problems before they become outages.
Let me walk you through building a system that understands event-driven patterns deeply.
The Challenge with Event-Driven Systems
When you move to event-driven architecture, you’re trading some operational simplicity for architectural flexibility. The tradeoff is worth it, but the complexity is real.
Consider this scenario: you have 30 microservices communicating through events. Service A publishes OrderCreated events. Services B, C, and D consume them. But are they all expecting the same schema? What if Service E added a field last month? Did everyone handle it gracefully? What happens if a message gets redelivered — did the consumer implement idempotency?
Now multiply that across hundreds of events, thousands of message producers and consumers, and you’ve got a sprawling system where subtle bugs hide. A customer’s order gets processed twice. A refund happens but the payment system doesn’t know. A workflow gets stuck in a dead letter queue and nobody notices for three days.
Claude Code solves this by understanding event schemas, routing patterns, failure handling, and consistency requirements. It can review your entire event architecture and spot issues that would take a human weeks to find.
The Hidden Layer: Why Event-Driven Complexity Matters
Before we build a solution, let’s understand why event-driven architecture creates so many problems. Traditional request-response systems have one clear path: A calls B, B returns a response, done. If B fails, A knows immediately and can handle the error.
Event-driven systems are different. A publishes an event and moves on. B might consume it immediately, or hours later. C might consume it and decide not to do anything. D might consume it and publish a new event, triggering F and G. The causality chains become complex. Debugging becomes a nightmare.
Here’s the pain: when things go wrong in an event-driven system, you have no idea where to look. An order gets processed twice? Was it the producer publishing twice, or the consumer processing twice, or a network retry that duplicated the message? A payment doesn’t correlate to an order? Did the order event fail to produce, or did the payment event get lost? Did the consumer crash before finishing its transaction?
Traditional monitoring and debugging tools don’t help. They show you individual services, not the event flow. They show you that a message was processed, not whether it was processed correctly. They don’t show you which events are orphaned in dead letter queues or why they ended up there.
This is where Claude Code becomes invaluable. It can read your event schemas, trace your producer and consumer code, review your retry logic, and understand your intent. It can then reason about correctness: “If OrderCreated publishes but PaymentAuthorized never arrives, the event flow breaks. Here’s the risk.”
Architecture Overview
We’re building a system with four main components:
- Schema analyzer — Validates event schemas for versioning and compatibility
- Flow mapper — Traces event paths from producers through consumers
- Failure auditor — Reviews retry logic, dead letters, and recovery patterns
- Intelligence engine — Uses Claude to reason about ordering, idempotency, and consistency
The key insight: event-driven architectures have patterns. Once Claude understands those patterns, it can spot violations automatically.
Why This Matters for Your Organization
The complexity of event-driven systems has direct business impact. In traditional systems, you catch errors before they ship. In event-driven systems, errors might not manifest for hours or days. A customer’s refund might get stuck because the consumer never restarted after a crash. A fraud detection event might arrive too late to stop a transaction.
Worse, event-driven systems hide these problems. The producer succeeded. The consumer ran successfully. From both perspectives, everything worked. But the overall system failed because some subtle ordering or timing issue caused a gap.
Event-driven architecture review isn’t theoretical. It’s practical risk management. You’re trying to identify failure modes before they cost you money, trust, or reputation.
Step 1: Define Event Schema Models
First, let’s create TypeScript interfaces that represent event definitions:
// src/models/event-schema.ts
export interface EventSchema {
name: string;
version: string;
namespace?: string;
description?: string;
fields: EventField[];
tags?: string[];
producedBy: string[]; // Service names that produce this
consumedBy: string[]; // Service names that consume this
deprecated?: boolean;
deprecatedAt?: Date;
successor?: string; // Name of the new event to use
createdAt: Date;
lastModifiedAt: Date;
}
export interface EventField {
name: string;
type: "string" | "number" | "boolean" | "object" | "array" | "timestamp";
required: boolean;
description?: string;
example?: any;
nested?: EventField[]; // For object types
itemType?: "string" | "number" | "boolean" | "object" | "timestamp"; // For array types
}
export interface EventMessage {
id: string; // Unique message ID for idempotency
timestamp: Date;
schemaName: string;
schemaVersion: string;
correlationId?: string; // Links related events
causationId?: string; // Links cause and effect
payload: Record<string, any>;
metadata?: {
source: string; // Service that produced it
retryCount?: number;
deadLettered?: boolean;
deadLetterReason?: string;
};
}
export interface ServiceDefinition {
name: string;
description?: string;
producesEvents: string[]; // Event names produced
consumesEvents: string[]; // Event names consumed
deathLetterHandler?: boolean; // Does it handle DLQ messages?
supportsIdempotency?: boolean;
orderingGuarantees?: "none" | "partitioned" | "strict";
}
export interface EventConsumerBinding {
consumerService: string;
eventSchema: string;
eventVersion: string;
handler: {
name: string;
timeout?: number; // milliseconds
retryPolicy?: RetryPolicy;
idempotencyKey?: string; // Which field is used for dedup
};
}
export interface RetryPolicy {
maxAttempts: number;
initialDelayMs: number;
maxDelayMs: number;
backoffMultiplier: number;
deadLetterAfterFail: boolean;
retryableErrorCodes?: string[];
}
export interface DeadLetterConfig {
queueName: string;
maxMessages?: number;
messageRetentionDays?: number;
alertThresholds?: {
criticalAfterMinutes?: number;
alertOnFirstMessage?: boolean;
};
reprocessingStrategy?: "manual" | "automatic" | "scheduled";
}
These interfaces represent the structure of event-driven systems. Real systems might have variations, but these capture the essentials.
Understanding Event Schema Evolution
One of the most common failure modes in event-driven systems is schema drift. Service A updates its OrderCreated event to include a “shippingAddress” field. Service A happily publishes the new version. But Service C never updates its consumer code to handle the new field. It keeps using the old schema definition. When it tries to validate incoming events, they fail validation because of the new required field.
Now you have a silent failure. The events are published. The consumer is running. But messages are silently dropped because they don’t match the expected schema. No error gets logged. No alert fires. The system looks healthy but is actually broken.
This is why schema versioning is critical. You need a clear contract: “This is version 2 of OrderCreated. It’s backward compatible with version 1 because it only added optional fields.” Then consumers can be updated on their own schedule, and you have time to verify before breaking changes.
Claude Code can analyze these contracts and spot violations. It can read your schema definitions and your consumer code, then reason about whether they match. It can tell you: “Service C is still using the old schema format. Upgrading to the new format will require these changes…”
The Consistency Challenge
Event-driven systems create consistency challenges that don’t exist in traditional systems. In a synchronous system, if a transaction fails, you know immediately and can roll back. In an event-driven system, you publish an event and move on. The consumer might fail to process it. Now you have two systems with inconsistent state.
How do you fix this? You need a way to detect inconsistency and recover. This might be a reconciliation process that runs periodically and compares state. It might be a compensating transaction that undoes the operation. It might be a retry mechanism that repeatedly attempts to process the failed event.
The hard part is deciding which approach fits your scenario. A payment failure needs immediate compensation. A logging event failure might just be retried. An analytics event failure might be acceptable to lose. Each requires different handling.
Claude Code can’t tell you which choice is right—that’s a business decision. But it can analyze your system and tell you: “This event flow doesn’t have a failure recovery mechanism. If the consumer crashes here, the transaction is orphaned.” Armed with that information, you can make a conscious decision about the risk.
The Idempotency Requirement
One concept that separates competent event-driven systems from broken ones is idempotency. A message might arrive once, or it might arrive multiple times due to retries, network duplicates, or reprocessing from dead letters.
If your consumer isn’t idempotent, processing the same message twice will cause problems. Transfer 100 dollars twice instead of once. Record two shipments instead of one. Deduct the same tax twice.
Idempotent processing is hard to get right. You typically use an idempotency key—a unique identifier from the event—and check before processing: “Have I already processed event ID X?” If yes, skip it. If no, process and record the ID.
The challenge is maintaining that idempotency key across failures. If your process crashes after processing but before writing the idempotency key, you’ll process the same event twice. You need transactional semantics: either the processing and the key recording both succeed, or both fail.
Claude Code can review your consumer code and check: “Is this consumer idempotent? How does it prevent duplicate processing? What happens if the recording of the idempotency key fails?” This is exactly the kind of semantic analysis that static tools can’t do.
Ordering Guarantees and Causality
Event-driven systems create assumptions about ordering. Service A publishes OrderCreated. Then OrderShipped. A consumer expects them in that order: a customer’s order is created, then shipped. If they arrive out of order, the consumer might try to ship an order that hasn’t been created yet.
Some message brokers provide strict ordering: if you publish OrderCreated then OrderShipped from a producer, they arrive in that order to consumers. Other brokers only provide partitioned ordering: messages with the same partition key stay ordered, but messages with different keys might arrive out of order.
This matters for your application’s correctness. If you rely on strict ordering but use a broker with only partitioned ordering, you have a latent bug. The system works most of the time, but under load or high volume, messages get reordered and cause failures.
Your event architecture review needs to understand these guarantees and check for violations. “This event flow depends on strict ordering but you’re using Kafka with unordered topics. Here’s the failure scenario…”
Dead Letter Queues: The Forgotten Responsibility
Every production event system needs dead letter queue handling. When a message fails processing (permanently, not just a transient error), it goes to a dead letter queue. From there, you need to investigate why it failed and handle it.
But many teams treat dead letter queues as a black hole. Messages arrive and nobody looks at them. Days later, a customer complains about a missing order or unprocessed refund. Investigation shows the message was in the dead letter queue for a week.
This is catastrophic because the failure was silent. The system looked healthy. The consumer was running. But messages were being lost.
Your event-driven system needs monitoring and alerting for dead letter queues. When a message lands in the DLQ, someone should be notified. The message should be investigated. A decision should be made: retry it, discard it, or escalate it.
Claude Code can analyze your event architecture and ask: “What happens to messages that fail processing? Do you have monitoring? What’s your SLA for responding to dead letter queue messages?” These are operational questions that affect reliability.
The Correlation Problem
In distributed systems, a customer’s action might trigger multiple events flowing through multiple services. OrderCreated triggers PaymentAuthorized, which triggers ShipmentScheduled, which triggers CarrierNotification. These events need to be correlated so you can ask: “What was the outcome of this customer’s request?”
Correlation is usually done with correlation IDs or causation IDs. OrderCreated generates a correlation ID. Every downstream event includes that ID. When you query for that ID, you see the entire flow.
But maintaining correlation IDs is tedious. Engineers forget to propagate them. Events get created without correlation IDs. Then when you try to investigate, the flow is fragmented.
Your event-driven system review should check: “Do all events have correlation IDs? Are they propagated consistently? Can I trace a customer request through your entire system?” Missing correlation makes debugging production issues exponentially harder.
Performance Considerations in Event-Driven Systems
Event-driven systems create different performance characteristics than traditional systems. A synchronous request-response system has latency measured in milliseconds. An event-driven system has latency measured in seconds (the time between publishing and consumer processing).
This latency needs to be acceptable for your use case. If you need real-time notification of an order status change, event-driven with seconds of latency might not work. If you can tolerate eventual consistency, it’s fine.
Your architecture review should assess: “Is the latency profile acceptable for this use case? Are you over-using asynchronous patterns where synchronous would be better?” Event-driven is powerful but not optimal for everything.
There are also throughput and backlog considerations. If a consumer is slow, events might pile up faster than they’re processed. You need monitoring to catch this before backlog becomes catastrophic. A consumer that processes 100 messages/second but is receiving 500 messages/second will eventually cause data loss or timeout failures.
Versioning and Evolution Strategy
One critical aspect of event-driven systems that teams often overlook until it’s too late is versioning strategy. How do you evolve event schemas without breaking existing consumers?
There are several approaches, each with tradeoffs. The simplest is strict backward compatibility: when you update a schema, you only add optional fields. Existing consumers can handle the new events because they simply ignore fields they don’t understand. When old consumers are updated, they start using the new fields.
But strict backward compatibility limits how much you can evolve. You can’t remove fields. You can’t change field types. Over time, your schema becomes bloated with deprecated fields that exist only for historical reasons.
Another approach is versioning: OrderCreated/v1, OrderCreated/v2, etc. Each version is independent. Consumers choose which version they consume. New services use the latest version. Old services keep using their version. You support multiple versions for a transition period, then deprecate old versions.
This works but requires more infrastructure. You need to handle multiple versions of the same event. You need migration tooling to help consumers upgrade.
A third approach is event upcasting: events are published in one version but transformed to a different version as they’re consumed. A consumer that understands version 2 can automatically upgrade version 1 events. This minimizes the number of versions you need to maintain.
Your event architecture review should examine: “What’s your versioning strategy? How do you handle schema evolution? How long do you support old versions? What happens when you need to break compatibility?”
Testing Event-Driven Architectures
Testing event-driven systems is harder than testing synchronous systems. You can’t just mock a function call—you need to simulate message publishing, delivery, and consumption, including failure modes.
A comprehensive test suite for event-driven systems should cover:
Happy path testing: Publish an event, verify the consumer processes it correctly. Verify that side effects happened. This is straightforward but only covers the sunny day scenario.
Schema validation testing: Publish events with missing required fields, extra fields, wrong types. Verify the consumer handles them gracefully (either rejects them clearly or handles forward/backward compatibility). This catches schema issues before they hit production.
Failure mode testing: Publish an event, then simulate consumer failure (crash, timeout, exception). Verify the event gets retried. Verify it eventually reaches the dead letter queue if max retries exceed. This is where most event-driven systems fail in testing.
Ordering testing: Publish events out of order and verify the system either handles it correctly or fails predictably. Publish events concurrently and verify no race conditions occur.
Idempotency testing: Publish the same event twice and verify it’s only processed once (or if not idempotent, verify why). This is critical and often overlooked.
End-to-end flow testing: Publish an event and trace it through all consumers and downstream events. Verify the entire flow completes correctly. This is where integration tests come in.
Many teams under-invest in these tests because they’re harder to write than unit tests. But they catch real failures that unit tests miss. Your event-driven architecture review should check: “What test coverage do you have for event flows? Do you test failure modes?”
Monitoring and Observability for Events
You can’t manage what you don’t measure. Event-driven systems need comprehensive monitoring of event flows, not just individual service metrics.
Key metrics to monitor:
Event publishing rates: How many events are being published per second? Are rates stable or spiking unexpectedly? A sudden spike might indicate a system malfunction.
Consumer lag: How far behind is a consumer compared to the producer? If lag is growing, the consumer is falling behind and backlog is accumulating.
Dead letter queue depth: How many messages are in the DLQ and how long have they been there? A growing DLQ indicates failures that need investigation.
End-to-end latency: Time from publishing an event to consumer processing completion. This should be consistent. Sudden spikes indicate problems.
Processing success rate: Percentage of messages processed successfully vs. failed. Even 99% success means 1% of messages are failing silently.
Message loss: In distributed systems, messages might be lost or duplicated. Monitoring should detect and alert on these.
Without this monitoring, failures hide. Your system looks healthy but is silently losing data. Customers report missing orders or unprocessed refunds. Investigation uncovers weeks of silent failures.
Your architecture review should verify: “Do you have monitoring for these metrics? What are your alerting thresholds? What’s your runbook for responding to alerts?”
Troubleshooting Event-Driven Failures
When event-driven systems fail, diagnosing the problem is exponentially harder than with synchronous systems. A customer’s order might be missing after 6 hours. Where did it get stuck? Did the OrderCreated event publish? Did PaymentAuthorized fail? Did ShipmentScheduled timeout?
To debug this, you need comprehensive tracing. Every event should include a correlation ID. Every hop should log with that ID. When you search for that correlation ID, you see the entire flow.
Without correlation IDs, you’re searching for a needle in a haystack. You know something failed, but you can’t trace the failure path. You’re left guessing which system to investigate.
Your event-driven architecture review should verify: “Does every event have a correlation ID? Are they logged? Can you reconstruct the flow from logs?” This is practical infrastructure that prevents debugging nightmares.
Another aspect is understanding event reprocessing. If a consumer failed and you reprocess the event from the dead letter queue, what happens? Is it idempotent? Does the correlation ID change? Does the system understand that this is a retry vs. a new event?
Unexpected reprocessing has caused data corruption in many systems. Your architecture should handle reprocessing explicitly: either prevent it or make all consumers handle it safely.
The Human Factor in Event-Driven Systems
Finally, there’s the human factor. Event-driven systems are cognitively harder for engineers to reason about than synchronous systems. The causality is less obvious. Debugging is harder. It’s easy to create subtle bugs that only appear under specific conditions.
This means your event-driven architecture needs good documentation, clear contracts, and strong typing (if using typed event definitions). It needs monitoring that makes failures visible. It needs patterns that engineers can understand and apply consistently.
Your architecture review should also assess team readiness. “Is your team comfortable with eventual consistency? Do they understand idempotency patterns? Have they experienced event-driven failures before?” Teams that understand the challenges are less likely to make mistakes.
Step 2: Schema Versioning and Compatibility Checker
Event schemas evolve over time. The first critical check is whether schema changes are backward compatible:
// src/analysis/schema-compatibility.ts
export interface CompatibilityIssue {
severity: "low" | "medium" | "high" | "breaking";
fieldName?: string;
issue: string;
explanation: string;
remediation: string;
}
export interface SchemaCompatibilityReport {
eventName: string;
fromVersion: string;
toVersion: string;
compatible: boolean;
issues: CompatibilityIssue[];
summary: string;
migratedConsumersCount?: number;
unmigratedConsumersCount?: number;
}
export class SchemaCompatibilityChecker {
private claude: Anthropic;
constructor(apiKey: string = process.env.ANTHROPIC_API_KEY || "") {
this.claude = new Anthropic({ apiKey });
}
async compareSchemas(
oldSchema: EventSchema,
newSchema: EventSchema,
): Promise<SchemaCompatibilityReport> {
// First, do basic structural checks
const basicIssues = this.checkBasicCompatibility(oldSchema, newSchema);
// Then ask Claude to do semantic analysis
const semanticIssues = await this.analyzeSemanticChanges(
oldSchema,
newSchema,
);
const allIssues = [...basicIssues, ...semanticIssues];
const breaking = allIssues.some((i) => i.severity === "breaking");
return {
eventName: oldSchema.name,
fromVersion: oldSchema.version,
toVersion: newSchema.version,
compatible: !breaking,
issues: allIssues,
summary: this.generateSummary(allIssues, breaking),
};
}
private checkBasicCompatibility(
oldSchema: EventSchema,
newSchema: EventSchema,
): CompatibilityIssue[] {
const issues: CompatibilityIssue[] = [];
// Check for required field removals (breaking change)
for (const oldField of oldSchema.fields) {
const newField = newSchema.fields.find((f) => f.name === oldField.name);
if (!newField && oldField.required) {
issues.push({
severity: "breaking",
fieldName: oldField.name,
issue: `Required field '${oldField.name}' was removed`,
explanation:
"Consumers expecting this field will receive messages without it, causing deserialization failures.",
remediation:
"Either restore the field or implement a migration period where the field is still included as optional.",
});
}
}
// Check for type changes on existing fields (usually breaking)
for (const newField of newSchema.fields) {
const oldField = oldSchema.fields.find((f) => f.name === newField.name);
if (oldField && oldField.type !== newField.type) {
issues.push({
severity: "breaking",
fieldName: newField.name,
issue: `Type of '${newField.name}' changed from ${oldField.type} to ${newField.type}`,
explanation:
"Consumers with strict type checking will fail to deserialize messages.",
remediation:
"Create a new event version or add a new field with the new type and keep the old one for backward compatibility.",
});
}
}
// Check for new required fields (breaking change)
for (const newField of newSchema.fields) {
const oldField = oldSchema.fields.find((f) => f.name === newField.name);
if (!oldField && newField.required) {
issues.push({
severity: "breaking",
fieldName: newField.name,
issue: `New required field '${newField.name}' added`,
explanation:
"Old consumers won't provide this field, causing validation failures.",
remediation:
"Make the field optional, provide a default value, or ensure all consumers are updated before enforcing.",
});
}
}
return issues;
}
private async analyzeSemanticChanges(
oldSchema: EventSchema,
newSchema: EventSchema,
): Promise<CompatibilityIssue[]> {
const prompt = `Review these two versions of an event schema and identify semantic compatibility issues:
Old Schema (v${oldSchema.version}):
${JSON.stringify(oldSchema.fields, null, 2)}
New Schema (v${newSchema.version}):
${JSON.stringify(newSchema.fields, null, 2)}
Look for:
1. Changes in field meaning or semantics
2. Enum value changes that could break state machines
3. New constraints or validations that old producers can't satisfy
4. Changes to field relationships or invariants
5. Changes that affect ordering or causality assumptions
Respond with JSON array of issues with: fieldName, issue, explanation, remediation, severity.`;
const message = await this.claude.messages.create({
model: "claude-3-5-sonnet-20241022",
max_tokens: 1024,
messages: [
{
role: "user",
content: prompt,
},
],
});
const responseText =
message.content[0].type === "text" ? message.content[0].text : "";
try {
const jsonMatch = responseText.match(/\[[\s\S]*\]/);
if (!jsonMatch) return [];
const parsed = JSON.parse(jsonMatch[0]);
return parsed.map((issue: any) => ({
severity: issue.severity || "medium",
fieldName: issue.fieldName,
issue: issue.issue,
explanation: issue.explanation,
remediation: issue.remediation,
}));
} catch (error) {
console.error("Failed to parse schema analysis:", error);
return [];
}
}
private generateSummary(
issues: CompatibilityIssue[],
breaking: boolean,
): string {
if (breaking) {
return `BREAKING CHANGES DETECTED. ${issues.length} issue(s) found. All consumers must be updated.`;
}
if (issues.length === 0) {
return "Fully backward compatible. Existing consumers will continue to work.";
}
const breaking_count = issues.filter(
(i) => i.severity === "breaking",
).length;
return `${breaking_count} potential issues. Review before deploying to production.`;
}
}
The key here is that we’re checking both structural changes (field removed, type changed) and semantic changes. Claude is particularly good at the latter — spotting that a change from "status": "pending|completed" to "status": "pending|processing|completed" adds a new state that downstream systems might not handle.
Step 3: Event Flow Mapping
Now let’s trace how events flow through your system:
// src/analysis/event-flow-mapper.ts
EventSchema,
ServiceDefinition,
EventConsumerBinding,
} from "../models/event-schema";
export interface EventFlow {
eventName: string;
producer: ServiceDefinition;
consumers: Array<{
service: ServiceDefinition;
binding: EventConsumerBinding;
handlerName: string;
}>;
cyclicDependencies?: string[];
orphanedConsumers?: string[]; // Consuming events that no one produces
orphanedProducers?: string[]; // Producing events that no one consumes
}
export interface SystemFlowGraph {
services: Map<string, ServiceDefinition>;
events: Map<string, EventSchema>;
bindings: EventConsumerBinding[];
flows: EventFlow[];
issues: FlowIssue[];
}
export interface FlowIssue {
type:
| "orphaned-event"
| "cyclic-dependency"
| "missing-handler"
| "version-mismatch"
| "ordering-violation"
| "idempotency-gap";
severity: "low" | "medium" | "high";
description: string;
affectedServices: string[];
recommendation: string;
}
export class EventFlowMapper {
mapEventFlows(
services: ServiceDefinition[],
schemas: EventSchema[],
bindings: EventConsumerBinding[],
): SystemFlowGraph {
const serviceMap = new Map(services.map((s) => [s.name, s]));
const schemaMap = new Map(
schemas.map((s) => [`${s.name}:${s.version}`, s]),
);
const flows: EventFlow[] = [];
const issues: FlowIssue[] = [];
// Build flows for each event
for (const schema of schemas) {
const producers = services.filter((s) =>
s.producesEvents.includes(schema.name),
);
if (producers.length === 0) {
issues.push({
type: "orphaned-event",
severity: "low",
description: `Event '${schema.name}' is defined but no service produces it`,
affectedServices: [],
recommendation: "Remove unused event schema or add a producer.",
});
continue;
}
for (const producer of producers) {
const consumers = services.filter((s) =>
s.consumesEvents.includes(schema.name),
);
if (consumers.length === 0) {
issues.push({
type: "orphaned-event",
severity: "low",
description: `Event '${schema.name}' produced by '${producer.name}' but has no consumers`,
affectedServices: [producer.name],
recommendation: "Remove the producer or add consumers.",
});
}
flows.push({
eventName: schema.name,
producer,
consumers: consumers.map((consumer) => {
const binding = bindings.find(
(b) =>
b.consumerService === consumer.name &&
b.eventSchema === schema.name,
);
return {
service: consumer,
binding: binding!,
handlerName: binding?.handler.name || "unknown",
};
}),
});
}
}
// Check for cyclic dependencies
const cyclics = this.detectCyclicDependencies(flows);
for (const cyclic of cyclics) {
issues.push({
type: "cyclic-dependency",
severity: "high",
description: `Cyclic event dependency detected: ${cyclic.join(" -> ")}`,
affectedServices: cyclic,
recommendation:
"Refactor to break the cycle, possibly by introducing a new service or event.",
});
}
return {
services: serviceMap,
events: schemaMap,
bindings,
flows,
issues,
};
}
private detectCyclicDependencies(flows: EventFlow[]): string[][] {
const graph = new Map<string, string[]>();
// Build adjacency list: service A -> service B if A produces an event B consumes
for (const flow of flows) {
const producer = flow.producer.name;
if (!graph.has(producer)) {
graph.set(producer, []);
}
for (const consumer of flow.consumers) {
graph.get(producer)!.push(consumer.service.name);
}
}
const visited = new Set<string>();
const cycles: string[][] = [];
for (const node of graph.keys()) {
if (!visited.has(node)) {
const path: string[] = [];
this.dfsCycle(node, graph, visited, path, cycles);
}
}
return cycles;
}
private dfsCycle(
node: string,
graph: Map<string, string[]>,
visited: Set<string>,
path: string[],
cycles: string[][],
): void {
visited.add(node);
path.push(node);
for (const neighbor of graph.get(node) || []) {
const cycleStart = path.indexOf(neighbor);
if (cycleStart !== -1) {
// Found a cycle
cycles.push(path.slice(cycleStart).concat(neighbor));
} else if (!visited.has(neighbor)) {
this.dfsCycle(neighbor, graph, visited, path, cycles);
}
}
path.pop();
}
}
This mapper builds a graph of your event flows and automatically detects cycles (where service A depends on events from B, which depends on A). These cycles usually indicate architectural problems.
Step 4: Idempotency and Ordering Analyzer
This is crucial. Event-driven systems have to handle message redelivery. If an event gets delivered twice, what happens?
// src/analysis/idempotency-analyzer.ts
export interface IdempotencyIssue {
consumerService: string;
eventSchema: string;
severity: "low" | "medium" | "high" | "critical";
issue: string;
riskDescription: string;
recommendation: string;
}
export interface IdempotencyAudit {
consumerService: string;
eventSchema: string;
isIdempotent: boolean;
idempotencyKey?: string;
detectDuplicateHow?: string; // "message-id" | "business-key" | "state-check" | "unknown"
issues: IdempotencyIssue[];
recommendations: string[];
}
export class IdempotencyAnalyzer {
private claude: Anthropic;
constructor(apiKey: string = process.env.ANTHROPIC_API_KEY || "") {
this.claude = new Anthropic({ apiKey });
}
async analyzeIdempotency(
binding: EventConsumerBinding,
eventSchema: EventSchema,
): Promise<IdempotencyAudit> {
const issues: IdempotencyIssue[] = [];
// If the binding explicitly says it doesn't handle idempotency, flag it
if (binding.handler.idempotencyKey === undefined) {
issues.push({
consumerService: binding.consumerService,
eventSchema: eventSchema.name,
severity: "critical",
issue: "No idempotency key configured",
riskDescription: `If '${eventSchema.name}' is redelivered (due to network failure, timeout, or retry), the handler will process it multiple times. This could lead to duplicate payments, double-charged customers, inventory inconsistencies, etc.`,
recommendation:
"Define an idempotency key field from the event (e.g., the event's unique ID) and implement deduplication using a cache or database.",
});
}
// Ask Claude to review the event schema and handler for semantic idempotency issues
const semanticIssues = await this.analyzeSemanticIdempotency(
binding,
eventSchema,
);
issues.push(...semanticIssues);
const isIdempotent =
issues.filter((i) => i.severity === "critical").length === 0;
const recommendations = [...new Set(issues.map((i) => i.recommendation))];
return {
consumerService: binding.consumerService,
eventSchema: eventSchema.name,
isIdempotent,
idempotencyKey: binding.handler.idempotencyKey,
detectDuplicateHow: this.inferIdempotencyStrategy(binding, eventSchema),
issues,
recommendations,
};
}
private async analyzeSemanticIdempotency(
binding: EventConsumerBinding,
eventSchema: EventSchema,
): Promise<IdempotencyIssue[]> {
const prompt = `You are reviewing event handler idempotency. A service '${binding.consumerService}'
consumes event '${eventSchema.name}' with handler '${binding.handler.name}'.
Event fields:
${JSON.stringify(eventSchema.fields, null, 2)}
Handler configuration:
- Idempotency key: ${binding.handler.idempotencyKey || "NOT CONFIGURED"}
- Max retries: ${binding.handler.retryPolicy?.maxAttempts || 3}
Identify potential issues:
1. State mutations that aren't idempotent (e.g., incrementing a counter)
2. Side effects that can't be safely repeated (e.g., sending emails, posting to external APIs)
3. Ordering assumptions that could be violated by redelivery
4. Business logic that assumes "exactly once" delivery
Respond with JSON array of issues with: issue, riskDescription, recommendation, severity.`;
const message = await this.claude.messages.create({
model: "claude-3-5-sonnet-20241022",
max_tokens: 1024,
messages: [
{
role: "user",
content: prompt,
},
],
});
const responseText =
message.content[0].type === "text" ? message.content[0].text : "";
try {
const jsonMatch = responseText.match(/\[[\s\S]*\]/);
if (!jsonMatch) return [];
const parsed = JSON.parse(jsonMatch[0]);
return parsed.map((item: any) => ({
consumerService: binding.consumerService,
eventSchema: eventSchema.name,
severity: item.severity || "medium",
issue: item.issue,
riskDescription: item.riskDescription,
recommendation: item.recommendation,
}));
} catch (error) {
console.error("Failed to parse idempotency analysis:", error);
return [];
}
}
private inferIdempotencyStrategy(
binding: EventConsumerBinding,
eventSchema: EventSchema,
): string {
const idempotencyKey = binding.handler.idempotencyKey;
if (!idempotencyKey) return "unknown";
// Check if they're using message ID
if (idempotencyKey === "id" || idempotencyKey === "messageId") {
return "message-id";
}
// Check if they're using a business key
const eventHasBusinessKey =
eventSchema.fields.some((f) => f.name === "orderId") ||
eventSchema.fields.some((f) => f.name === "transactionId") ||
eventSchema.fields.some((f) => f.name === "paymentId");
if (eventHasBusinessKey) {
return "business-key";
}
// Check if they might be doing state checks
if (binding.handler.timeout && binding.handler.timeout > 1000) {
return "state-check"; // Longer timeout suggests checking existing state
}
return "unknown";
}
}
Step 5: Dead Letter Queue Audit
Dead letter queues are where messages go to die. You need visibility into them:
// src/analysis/dlq-auditor.ts
export interface DLQAudit {
queueName: string;
messageCount: number;
oldestMessageAgeMinutes: number;
youngestMessageAgeMinutes: number;
issues: DLQIssue[];
recommendations: string[];
isHealthy: boolean;
}
export interface DLQIssue {
severity: "low" | "medium" | "high" | "critical";
type: string;
description: string;
suggestedAction: string;
}
export class DLQAuditor {
private claude: Anthropic;
constructor(apiKey: string = process.env.ANTHROPIC_API_KEY || "") {
this.claude = new Anthropic({ apiKey });
}
async auditDLQ(
config: DeadLetterConfig,
recentMessages: any[],
timeDataAvailable: number, // hours of monitoring data
): Promise<DLQAudit> {
const issues: DLQIssue[] = [];
// Check 1: Is the DLQ growing?
if (recentMessages.length > (config.maxMessages || 100)) {
issues.push({
severity: "high",
type: "DLQ_OVERFLOW",
description: `DLQ has ${recentMessages.length} messages (configured max: ${config.maxMessages})`,
suggestedAction:
"Increase max retention or investigate why messages are being dead-lettered.",
});
}
// Check 2: Are messages stuck in the DLQ?
const oldestMessage = recentMessages.length > 0 ? recentMessages[0] : null;
const messageAgeMinutes = oldestMessage
? Math.floor(
(Date.now() - new Date(oldestMessage.timestamp).getTime()) / 60000,
)
: 0;
const alertThreshold = config.alertThresholds?.criticalAfterMinutes || 60;
if (messageAgeMinutes > alertThreshold) {
issues.push({
severity: "critical",
type: "STALE_MESSAGES",
description: `Oldest DLQ message is ${messageAgeMinutes} minutes old (alert threshold: ${alertThreshold} minutes)`,
suggestedAction:
"Investigate the consumer failures. Manually reprocess or contact the team.",
});
}
// Check 3: Analyze error patterns with Claude
const errorAnalysis = await this.analyzeErrorPatterns(
recentMessages,
config,
);
issues.push(...errorAnalysis);
// Check 4: Review reprocessing strategy
if (
!config.reprocessingStrategy ||
config.reprocessingStrategy === "manual"
) {
issues.push({
severity: "medium",
type: "MANUAL_REPROCESSING",
description:
"DLQ uses manual reprocessing. Messages won't be retried automatically.",
suggestedAction:
"Consider implementing automatic or scheduled reprocessing for recoverable errors.",
});
}
const recommendations = this.generateRecommendations(issues);
const isHealthy =
issues.filter((i) => i.severity === "critical").length === 0;
return {
queueName: config.queueName,
messageCount: recentMessages.length,
oldestMessageAgeMinutes: messageAgeMinutes,
youngestMessageAgeMinutes:
recentMessages.length > 0
? Math.floor(
(Date.now() -
new Date(
recentMessages[recentMessages.length - 1].timestamp,
).getTime()) /
60000,
)
: 0,
issues,
recommendations,
isHealthy,
};
}
private async analyzeErrorPatterns(
messages: any[],
config: DeadLetterConfig,
): Promise<DLQIssue[]> {
if (messages.length === 0) return [];
// Group by error reason
const errorCounts: Record<string, number> = {};
const sampleErrors: Record<string, string> = {};
for (const msg of messages.slice(0, 50)) {
const reason = msg.deadLetterReason || "Unknown";
errorCounts[reason] = (errorCounts[reason] || 0) + 1;
if (!sampleErrors[reason]) {
sampleErrors[reason] = JSON.stringify(msg, null, 2).slice(0, 500);
}
}
const prompt = `Analyze these dead letter queue error patterns:
${Object.entries(errorCounts)
.map(
([reason, count]) =>
`- ${reason}: ${count} occurrences\n Sample: ${sampleErrors[reason]}`,
)
.join("\n")}
Queue config:
${JSON.stringify(config, null, 2)}
For each error pattern, identify:
1. Root cause
2. Whether it's transient or permanent
3. Recovery strategy
Respond with JSON array of issues: { type, description, suggestedAction, severity }.`;
const message = await this.claude.messages.create({
model: "claude-3-5-sonnet-20241022",
max_tokens: 1024,
messages: [
{
role: "user",
content: prompt,
},
],
});
const responseText =
message.content[0].type === "text" ? message.content[0].text : "";
try {
const jsonMatch = responseText.match(/\[[\s\S]*\]/);
if (!jsonMatch) return [];
const parsed = JSON.parse(jsonMatch[0]);
return parsed.map((item: any) => ({
severity: item.severity || "medium",
type: item.type,
description: item.description,
suggestedAction: item.suggestedAction,
}));
} catch (error) {
console.error("Failed to parse DLQ analysis:", error);
return [];
}
}
private generateRecommendations(issues: DLQIssue[]): string[] {
const recommendations = new Set<string>();
for (const issue of issues) {
if (issue.severity === "critical") {
recommendations.add("URGENT: " + issue.suggestedAction);
} else {
recommendations.add(issue.suggestedAction);
}
}
return Array.from(recommendations);
}
}
Step 6: Comprehensive Architecture Review
Now let’s tie everything together into a full architecture review:
// src/event-architecture-reviewer.ts
EventSchema,
ServiceDefinition,
EventConsumerBinding,
DeadLetterConfig,
} from "./models/event-schema";
export interface EventArchitectureReview {
reviewDate: Date;
systemName: string;
serviceCount: number;
eventCount: number;
overallHealth: "excellent" | "good" | "warning" | "critical";
sections: ReviewSection[];
executiveSummary: string;
actionItems: ActionItem[];
}
export interface ReviewSection {
title: string;
findings: string;
issues: string[];
recommendations: string[];
}
export interface ActionItem {
priority: "critical" | "high" | "medium" | "low";
description: string;
estimatedEffort?: string;
team?: string;
}
export class EventArchitectureReviewer {
private claude: Anthropic;
private compatibilityChecker: SchemaCompatibilityChecker;
private flowMapper: EventFlowMapper;
private idempotencyAnalyzer: IdempotencyAnalyzer;
private dlqAuditor: DLQAuditor;
constructor(apiKey: string = process.env.ANTHROPIC_API_KEY || "") {
this.claude = new Anthropic({ apiKey });
this.compatibilityChecker = new SchemaCompatibilityChecker(apiKey);
this.flowMapper = new EventFlowMapper();
this.idempotencyAnalyzer = new IdempotencyAnalyzer(apiKey);
this.dlqAuditor = new DLQAuditor(apiKey);
}
async reviewArchitecture(
systemName: string,
services: ServiceDefinition[],
schemas: EventSchema[],
bindings: EventConsumerBinding[],
dlqConfigs: DeadLetterConfig[],
dlqMessages: Map<string, any[]>, // Map of DLQ name -> sample messages
): Promise<EventArchitectureReview> {
const sections: ReviewSection[] = [];
// 1. Schema compatibility review
console.log("Analyzing schema compatibility...");
const schemaSection = await this.reviewSchemaCompatibility(schemas);
sections.push(schemaSection);
// 2. Event flow analysis
console.log("Mapping event flows...");
const flowSection = this.reviewEventFlows(services, schemas, bindings);
sections.push(flowSection);
// 3. Idempotency audit
console.log("Auditing idempotency...");
const idempotencySection = await this.reviewIdempotency(bindings, schemas);
sections.push(idempotencySection);
// 4. Dead letter queue review
console.log("Auditing dead letter queues...");
const dlqSection = await this.reviewDeadLetterQueues(
dlqConfigs,
dlqMessages,
);
sections.push(dlqSection);
// 5. Comprehensive assessment with Claude
console.log("Generating comprehensive assessment...");
const assessmentSection = await this.generateComprehensiveAssessment(
systemName,
services,
schemas,
sections,
);
sections.push(assessmentSection);
// Calculate overall health
const criticalIssues = sections.reduce(
(sum, s) => sum + s.issues.length,
0,
);
const overallHealth =
criticalIssues === 0
? "excellent"
: criticalIssues <= 2
? "good"
: criticalIssues <= 5
? "warning"
: "critical";
// Extract action items
const actionItems = this.extractActionItems(sections);
// Generate executive summary
const executiveSummary = this.generateExecutiveSummary(
systemName,
services.length,
schemas.length,
criticalIssues,
overallHealth,
);
return {
reviewDate: new Date(),
systemName,
serviceCount: services.length,
eventCount: schemas.length,
overallHealth,
sections,
executiveSummary,
actionItems,
};
}
private async reviewSchemaCompatibility(
schemas: EventSchema[],
): Promise<ReviewSection> {
// Find schema version pairs and check compatibility
const schemasByName = new Map<string, EventSchema[]>();
for (const schema of schemas) {
if (!schemasByName.has(schema.name)) {
schemasByName.set(schema.name, []);
}
schemasByName.get(schema.name)!.push(schema);
}
const issues: string[] = [];
const recommendations: string[] = [];
for (const [eventName, versions] of schemasByName) {
versions.sort((a, b) => a.version.localeCompare(b.version));
// Check each consecutive version pair
for (let i = 1; i < versions.length; i++) {
const report = await this.compatibilityChecker.compareSchemas(
versions[i - 1],
versions[i],
);
if (!report.compatible) {
issues.push(
`${eventName}: v${report.fromVersion} -> v${report.toVersion} has breaking changes`,
);
}
recommendations.push(...report.issues.map((iss) => iss.remediation));
}
}
return {
title: "Schema Versioning & Compatibility",
findings: `Analyzed ${schemas.length} event schemas across ${schemasByName.size} event types.`,
issues,
recommendations: [...new Set(recommendations)],
};
}
private reviewEventFlows(
services: ServiceDefinition[],
schemas: EventSchema[],
bindings: EventConsumerBinding[],
): ReviewSection {
const graph = this.flowMapper.mapEventFlows(services, schemas, bindings);
const issues = graph.issues.map((i) => i.description);
const recommendations = graph.issues.map((i) => i.recommendation);
return {
title: "Event Flow & Dependencies",
findings: `Mapped ${graph.flows.length} event flows across ${services.length} services.`,
issues,
recommendations: [...new Set(recommendations)],
};
}
private async reviewIdempotency(
bindings: EventConsumerBinding[],
schemas: EventSchema[],
): Promise<ReviewSection> {
const issues: string[] = [];
const recommendations: string[] = [];
for (const binding of bindings) {
const schema = schemas.find((s) => s.name === binding.eventSchema);
if (!schema) continue;
const audit = await this.idempotencyAnalyzer.analyzeIdempotency(
binding,
schema,
);
if (!audit.isIdempotent) {
issues.push(
`${binding.consumerService} -> ${binding.eventSchema}: Not idempotent`,
);
}
recommendations.push(...audit.recommendations);
}
return {
title: "Idempotency & Deduplication",
findings: `Reviewed idempotency for ${bindings.length} consumer-event bindings.`,
issues,
recommendations: [...new Set(recommendations)],
};
}
private async reviewDeadLetterQueues(
dlqConfigs: DeadLetterConfig[],
dlqMessages: Map<string, any[]>,
): Promise<ReviewSection> {
const issues: string[] = [];
const recommendations: string[] = [];
for (const config of dlqConfigs) {
const messages = dlqMessages.get(config.queueName) || [];
const audit = await this.dlqAuditor.auditDLQ(config, messages, 24);
if (!audit.isHealthy) {
issues.push(
`${config.queueName}: ${audit.issues.length} issues detected`,
);
}
recommendations.push(...audit.recommendations);
}
return {
title: "Dead Letter Queues",
findings: `Audited ${dlqConfigs.length} dead letter queues.`,
issues,
recommendations: [...new Set(recommendations)],
};
}
private async generateComprehensiveAssessment(
systemName: string,
services: ServiceDefinition[],
schemas: EventSchema[],
sections: ReviewSection[],
): Promise<ReviewSection> {
const prompt = `You are an expert in event-driven architecture. Review this system assessment:
System: ${systemName}
Services: ${services.length}
Event Types: ${schemas.length}
Review findings:
${sections.map((s) => `## ${s.title}\nIssues: ${s.issues.join(", ") || "None"}`).join("\n\n")}
Provide a comprehensive assessment that:
1. Identifies the top 3 architectural risks
2. Evaluates maturity (immature/developing/mature/advanced)
3. Suggests the most impactful improvements
4. Estimates the effort and ROI of key changes
Format as: MATURITY: [level], RISKS: [list], IMPROVEMENTS: [list], ROI: [assessment]`;
const message = await this.claude.messages.create({
model: "claude-3-5-sonnet-20241022",
max_tokens: 1024,
messages: [
{
role: "user",
content: prompt,
},
],
});
const responseText =
message.content[0].type === "text" ? message.content[0].text : "";
return {
title: "Comprehensive Assessment & Maturity",
findings: responseText,
issues: [],
recommendations: [],
};
}
private extractActionItems(sections: ReviewSection[]): ActionItem[] {
const actionItems: ActionItem[] = [];
for (const section of sections) {
for (const rec of section.recommendations) {
if (rec.includes("URGENT") || rec.includes("critical")) {
actionItems.push({
priority: "critical",
description: rec.replace("URGENT: ", ""),
team: section.title.includes("Schema") ? "Platform" : "Engineering",
});
} else if (rec.includes("high")) {
actionItems.push({
priority: "high",
description: rec,
team: "Engineering",
});
} else {
actionItems.push({
priority: "medium",
description: rec,
team: "Engineering",
});
}
}
}
return actionItems.sort((a, b) => {
const priorityOrder = { critical: 0, high: 1, medium: 2, low: 3 };
return priorityOrder[a.priority] - priorityOrder[b.priority];
});
}
private generateExecutiveSummary(
systemName: string,
serviceCount: number,
eventCount: number,
issueCount: number,
health: string,
): string {
return `
# Executive Summary
**System**: ${systemName}
**Status**: ${health.toUpperCase()}
**Services**: ${serviceCount}
**Event Types**: ${eventCount}
**Issues Found**: ${issueCount}
This event-driven architecture is operating at a ${health} level. Key focus areas include schema compatibility, idempotency guarantees, and dead letter queue handling. See action items for prioritized remediation.
`;
}
}
// Example usage
async function main() {
const reviewer = new EventArchitectureReviewer();
const services: ServiceDefinition[] = [
{
name: "payment-service",
description: "Handles payment processing",
producesEvents: ["PaymentProcessed", "PaymentFailed"],
consumesEvents: ["OrderCreated"],
supportsIdempotency: true,
orderingGuarantees: "partitioned",
},
{
name: "inventory-service",
description: "Manages inventory",
producesEvents: ["InventoryReserved"],
consumesEvents: ["OrderCreated", "PaymentProcessed"],
supportsIdempotency: true,
orderingGuarantees: "partitioned",
},
];
const schemas: EventSchema[] = [
{
name: "OrderCreated",
version: "1.0",
description: "Order has been created",
fields: [
{ name: "orderId", type: "string", required: true },
{ name: "customerId", type: "string", required: true },
{ name: "amount", type: "number", required: true },
],
producedBy: ["order-service"],
consumedBy: ["payment-service", "inventory-service"],
createdAt: new Date("2024-01-01"),
lastModifiedAt: new Date("2024-01-01"),
},
];
const bindings: EventConsumerBinding[] = [
{
consumerService: "payment-service",
eventSchema: "OrderCreated",
eventVersion: "1.0",
handler: {
name: "processPayment",
timeout: 5000,
idempotencyKey: "orderId",
},
},
];
const dlqConfigs: DeadLetterConfig[] = [
{
queueName: "payment-service-dlq",
maxMessages: 100,
messageRetentionDays: 30,
reprocessingStrategy: "manual",
},
];
const review = await reviewer.reviewArchitecture(
"E-Commerce Platform",
services,
schemas,
bindings,
dlqConfigs,
new Map(),
);
console.log("\n=== EVENT ARCHITECTURE REVIEW ===");
console.log(review.executiveSummary);
console.log("\n=== CRITICAL ACTION ITEMS ===");
review.actionItems
.filter((a) => a.priority === "critical")
.forEach((a) => console.log(`- ${a.description}`));
console.log("\n=== DETAILED FINDINGS ===");
review.sections.forEach((s) => {
console.log(`\n## ${s.title}`);
console.log(s.findings);
if (s.issues.length > 0) {
console.log("Issues:", s.issues);
}
});
}
main().catch(console.error);
Running this would output something like:
=== EVENT ARCHITECTURE REVIEW ===
# Executive Summary
**System**: E-Commerce Platform
**Status**: WARNING
**Services**: 2
**Event Types**: 1
**Issues Found**: 4
This event-driven architecture is operating at a warning level. Key focus areas include schema compatibility, idempotency guarantees, and dead letter queue handling. See action items for prioritized remediation.
=== CRITICAL ACTION ITEMS ===
- payment-service -> OrderCreated: Not idempotent - If payment events are redelivered, duplicate charges will occur
- DLQ has 47 messages - investigate processing failures
- Manual DLQ reprocessing increases MTTR - implement automatic retry for transient failures
=== DETAILED FINDINGS ===
## Schema Versioning & Compatibility
Analyzed 1 event schemas across 1 event types.
Issues: None
## Event Flow & Dependencies
Mapped 2 event flows across 2 services.
## Idempotency & Deduplication
Reviewed idempotency for 1 consumer-event bindings.
Issues: payment-service -> OrderCreated: Not idempotent
Why This Approach Works
Event-driven systems are fundamentally about patterns. Once Claude understands those patterns — schema evolution, idempotency, ordering, failure handling — it can apply them systematically across your entire architecture.
The key is that we’re not trying to understand every custom business logic. We’re understanding the infrastructure patterns and validating them.
Extending This System
You can add more analyzers for:
- Ordering guarantees: Do consumers assume strict ordering? Do producers guarantee it?
- Correlation tracking: Are causationId and correlationId used consistently?
- Latency profiles: Which event paths are latency-sensitive?
- Compliance checks: Are sensitive data fields encrypted in transit?
The real power comes when you run this regularly — monthly, or after each major deployment. You’ll catch drift before it becomes a production incident.
Implementing Event Architecture Review in Your Organization
Adopting a comprehensive event architecture review system requires both technical setup and organizational change. You need to document your event flows, make them machine-readable, and commit to regular analysis.
Start with a focused set of critical events—the ones that, if they fail, cause customer impact. OrderCreated, PaymentAuthorized, ShipmentNotified. Document each event: who produces it, who consumes it, what guarantees are needed, what failures are acceptable.
As you document events, you’ll discover gaps in your understanding. “Does Service C actually consume this event?” “What happens if the event arrives out of order?” These questions are valuable—they reveal architectural assumptions that need to be explicit.
Once events are documented, the architecture review becomes a tool for validating assumptions. Every few months, run the analysis and see what’s changed. Services that publish events now also consume them. Consumers have new handlers. Failure handling got more robust (or became riskier). The review captures these changes and flags issues.
Over time, this creates a virtuous cycle: Better documentation enables better analysis, which reveals improvement opportunities, which drive better architecture. Your event-driven system becomes more robust with each iteration.
This is event-driven architecture review in 2026 — automated, intelligent, and catching the subtle issues that slip past code reviews.
-iNet