Before the Clock Starts

Implementing Solace File Mesh for market infrastructure: What building disaster recovery replication for market infrastructure taught us, and how Solace File Mesh answers it.

Iesoft Technologies is a Solace partner. File Mesh Solution, Includes File Micro Integration & Solution Manager for “Design, Deploy & Govern”. Iesoft Technologies implements and operates it for customers. This is what we have learned doing so where the regulator is watching.

In 2021, the regulator of one of the world's busiest securities markets rewrote its recovery clock.

Market infrastructure institutions, the exchanges, clearing corporations and depositories that a market actually runs on, had until then been allowed several hours to recover from a disaster, and a generous share of that simply to decide that one had occurred. The revised framework cut the whole thing to well under an hour: minutes to declare, and a fixed number of minutes from that declaration to be running live from the disaster recovery site. Tolerable data loss was halved. Institutions had a single quarter to comply.

It was one regulator in one market, but the direction was not local. Operational resilience regimes in Europe, Southeast Asia and the Gulf have been compressing the same windows, for the same reason: a market that is down is a market failing at the one thing it exists to do.

Most of the commentary at the time was about the switchover. That is the visible part: the moment someone senior decides, and traffic moves to another building hundreds of kilometres away.

But the recovery window is not when you copy the data. By the time that clock starts, the data has to already be there. What the rule actually demands is that every file your risk, collateral, clearing and settlement systems depend on has been arriving at a second data centre continuously, correctly, all day, on a day when nothing was wrong and nobody was watching. The recovery window is not the replication window.

It is when you find out whether the replication was any good.

We implement that layer. As a Solace partner, we deploy and operate Solace File Mesh for institutions that live under this clock. What follows is what the clock does to an engineering problem, what is inside Solace's platform that answers it, and what we have learned putting it into production for our customers.


Part one: the brief the clock writes

Under the old framework, one clause quietly permitted an enormous simplification. Updates at the primary data centre had to be reflected at the recovery site immediately, which the same sentence then qualified as before end of day.

Before end of day is a batch window. A batch window means a scheduled job, a maintenance slot, and a comfortable assumption that if tonight's run fails you have until tomorrow morning. An entire generation of replication architecture was built on that sentence, and all of it was compliant.

A window measured in minutes, end to end, is a different problem.

The shape

Solace's Files Micro-Integration replicates files from a Source to a Sink, Solace's term for the receiving end, and the two halves never speak to each other directly. The Source sits beside the source filesystem, reads files, breaks them into parts and publishes each part as a guaranteed message onto the Solace event mesh. The Sink at the far site consumes those parts from a durable queue and writes files back out onto the destination filesystem.

Putting an event broker in the middle of a file transfer looks like an odd choice until you consider what the alternative asks you to build. A point-to-point protocol between two data centres gives you one path, and every property you care about, persistence, retry, ordering, delivery guarantees, multi-site fan-out, becomes your problem to implement and your problem to prove. Publishing onto a mesh that already provides those properties means the replication logic can be about files rather than about networks. The same stream reaches a near site and a disaster recovery site without the source knowing either exists.

It also means the two ends are decoupled in time. The Sink can be down for maintenance while the Source continues to publish. Messages wait in the queue. Nothing is lost and nothing needs coordinating.

Figure 1. The shape. The Source and Sink never speak directly. Data, command-centre events and the resume checkpoint all live in Solace Event Broker, and one publish reaches every site that listens.

The problem that actually matters

Here is the scenario that shapes everything else.

It is 14:20 on a trading day. A four gigabyte collateral file is 3.8 gigabytes of the way across. The WAN link drops.

The naive answer is to start again. Under an objective measured in hours, starting again is survivable. Under one measured in minutes it is not, and on a busy afternoon it may not even be survivable in absolute terms, because starting again means re-sending 3.8 gigabytes you have already sent, over a link that has just proven it can fail.

So the Source checkpoints. Before each part goes out it records where it has reached: the file identity, the total parts, the parts sent, and the exact byte read position. On restart it fetches that checkpoint, compares the file it finds against the file the checkpoint describes, seeks to the recorded offset and continues.

The interesting decision is where the checkpoint lives. It is a Last Value Queue, a Solace Event Broker construct that retains only the most recent message on a topic. The resume state sits in the message layer itself.

Consider what that avoids. No side database to provision, secure, back up and make highly available in both data centres. No local state file that can diverge from what the broker believes. No additional piece of infrastructure that has to be up for replication to be recoverable, and therefore no additional thing that can be down at 14:20. The broker is already the most available component in the architecture, because everything else depends on it. Putting the resume pointer inside it means the checkpoint has exactly the same availability as the transport.

Most teams treat a broker as a pipe. Solace's design treats it as a place to keep things. Fifteen years of building on Solace and Kafka have taught us how rare, and how right, that instinct is.


Figure 2. 14:20. The cumulative byte count travels on every part. The checkpoint lives in a Last Value Queue on Solace Event Broker, so it has the same availability as the transport.

Trusting nothing on the receiving side

Guaranteed messaging can redeliver. That is a feature, and it is why nothing gets lost when a consumer dies mid-processing. It also means the Sink must assume that any part may arrive twice, or that a part it has already written may show up again after a reconnect.

The obvious defence is to track part numbers. The Sink does not, because part numbers describe what the sender intended, and after a resume the sender's intentions and the receiver's disk may disagree.

Instead the Sync works in bytes. It derives the incoming part's start position from the cumulative byte count and the payload length, then compares that against how far the file on disk has actually got. Four outcomes:

Contiguous. The part begins exactly where the file ends. Write it.

Overlapping. The part begins before the current end of file but extends past it. This is what a resume looks like when the Source's checkpoint is slightly behind the Sink's disk. The buffer is trimmed to the new tail and only the genuinely new bytes are written.

Gapped. The part begins after the current end of file, meaning something in between never arrived. The part is not written, because writing it would produce a file that is the right length and wrong in the middle. That is the failure mode worth fearing: a corrupt file that looks fine.

Duplicate. The part covers ground already written. Discarded.


Figure 3. Four outcomes. Every incoming part is placed against the high-water mark of the file on disk. A gap is never written.

The overlapping case is the one that repays the effort. A resume across a WAN under load is not clean. Somebody's checkpoint is always a little stale. Handling that at byte granularity rather than part granularity is what lets a transfer resume without either losing a fragment or writing one twice.

Fidelity is not the same as bytes

A replica with the right contents and the wrong ownership is not a replica. When the recovery site has to run live operations, files have to be readable by the processes that expect to read them.

So the source's ownership and permissions travel with the data and are reapplied at the destination. Ownership is carried as user and group names rather than numeric ids, and resolved locally at the far end, because the same account can hold different numeric ids on different hosts and the numeric id is the thing that would silently be wrong.

Two payload types get their own handling. A zero-byte file is sent explicitly as an empty file rather than as a part with nothing in it, so the Sink cannot mistake it for a truncated transfer. A directory is its own message type and is created as a directory, so that empty directory structure survives replication. Neither is glamorous. Both exist because directory trees replicated without them come out subtly wrong, and subtly wrong is the expensive kind.

Four ways to replicate, and why one of them matters more now

The Files Micro-Integration handles a static file, replicated once per run. A directory, single level. A directory, recursive to arbitrary depth.

And a dynamic file, which is the one the regulation has been pushing toward. Here the Source replicates to end of file and then does not stop. It keeps watching, and as the source file grows it continues to publish. The transfer is bounded by the trading day rather than by the file's current length.

That mode exists because some files are written throughout the session rather than produced at the end of it. It has become more important since, because the same regulator has more recently moved the Recovery Point Objective from a matter of minutes to near zero, with zero data loss required for clearing corporations and depositories. A few minutes of tolerable loss can be met with scheduled replication. Near zero cannot. Continuous replication of a file that is still being written stops being a special case and starts being the default.

The details that come from having been paged

Some behaviour in the Files Micro-Integration only makes sense once you have operated it through something going wrong at an unhelpful hour. These are the details we point customers at first.

The Source's directory scan looks back a bounded number of days for modified files, four by default. That bound is there because of long weekends. Without it, a holiday during which someone touched an archive directory produces a mass re-replication on the next working morning, consuming exactly the bandwidth the morning needs.

There is a skip window that ignores files created within a configured interval of the last successful run, which is how you avoid picking up a file that is still being written by something else.

There are per-directory file ceilings, duplicate path detection across configuration entries, and wildcard whitelists and blacklists, because the scope of a replication is a thing that quietly grows.

There is bandwidth limiting, because a replication that saturates the inter-site link during market hours has solved one problem by creating a worse one.

And the scheduler refuses to start an instance that is already running. It checks, finds the process, logs, and exits rather than starting a second copy. Before it launches anything at all it opens a session to the event broker to confirm the user account is enabled, so that a permissions problem surfaces as a clear pre-flight failure rather than as a replication that appears to start and then quietly does nothing.

Two streams, from the beginning

The last thing about the micro-integration is the one that turned out to matter most.

File data goes to one topic. Everything the Source and Sink have to say about themselves goes to others: a start event when a run begins, heartbeats at a configured interval, warnings, errors, a multipart event, a completion event carrying the run's totals.

Nothing polls them. They publish.

It reads as a design convenience. In practice it is the whole architecture. The data path and the story of the data path are separate by design, which means the monitoring layer can be built, changed, and taken down without touching replication. It also meant that adding a new consumer of operational events required no change to the Source or Sink at all.

Part two: what one deployment cannot teach an industry

Everything above works. We have run it in production across primary, near-site and recovery data centres, on the feeds any market regulator would name as critical: risk, collateral, clearing and settlement.

Deployed on its own, without a management layer above it, it is also a system you have to be initiated into.

The configuration lives in files, one per instance, in a directory tree with a shape you learn by being shown. Topic strings are composed by hand and have to agree at both ends. Broker client profiles and access control profiles are created by hand, one per instance, granting each exactly the topics it needs. Scheduler entries are shell scripts. Job orchestration is crontab or Control-M depending on the site. Liveness is a process check every minute.

For one institution with a team that lives inside it, this is correct. Nothing in it is wasted, and the concreteness is a virtue when something breaks at 06:40.

But the clock is not one institution's problem. It landed on every exchange, clearing corporation and depository in that market at once, with one quarter to comply. It was later extended to other systemically important participants, on the explicit reasoning that a registrar holding tens of millions of investor accounts is infrastructure too. And then the target moved again, to near zero data loss, with a new requirement for a documented reconciliation process when operations resume at another site.

You cannot hand an industry a tarball and a set of conventions. Not because the engineering does not transfer, but because the knowledge does not.

That question, how do you give this to the hundredth institution without also giving them the team, is what Solace File Mesh Manager exists to answer.

Part three: Solace File Mesh Manager

Solution Manager

Solution Manager is where the hand-edited configuration file goes to die.

A data flow becomes a first-class object: named, versioned, and built visually from a catalogue of sources, destinations and transformations. That much is table stakes and every integration product claims it. What matters is the rigour around it.

Connections are objects in their own right, reusable across flows, and they are tested before they can be saved. A connection that has never successfully connected does not become part of a flow. Once connected, the designer browses the actual schemas, tables and columns on the far side, so field mapping happens against the system as it exists rather than against documentation describing how it existed.

The database coverage tells you what kind of estate Solace built this for. PostgreSQL, Oracle, MySQL and MariaDB, SQL Server, SAP HANA, Vertica. Nobody accumulates that list building for greenfield.

Flows move in both directions: database to event mesh, and event mesh to database. Real estates need both, usually on the same afternoon.

From designed to running

There is a gap in most integration platforms where the product stops and a runbook begins. You design something, the tool congratulates you, and then a human copies artefacts to a server.

File Mesh Manager closes it. A designed flow is packaged into a deployable connector, pushed to a Git repository, and deployed over SSH to a target server with image and version tag validation. The port it will use is registered against the organisation that owns it, so two connectors on the same host cannot collide. The process is started, and its status is visible from the same console that designed it.

Promotion between environments works the same way. A flow moves from development to test to production using the target environment's own approved connections, never the source environment's, and every promotion is recorded as a transaction.

The effect is that integration configuration behaves like code. Reviewable, diffable, revertable, attributable. Git is not a storage detail here. It is the system of record for what is deployed where and who put it there.

The fan-out

Behind File Mesh Manager's Command Center dashboard sit five listener services, and they are what the Source and Sink's event stream feeds.

A single Solace event stream from the running connectors fans out to five independent consumers, each owning exactly one concern. A state event listener writes the full file-transfer lifecycle history, routing malformed payloads to an error table rather than discarding them. A file complete listener updates per-connector statistics and triggers completion notifications. An error listener consumes error events and sends alerts throttled and de-duplicated against the run identifier it cached when the run began, so that one failing run produces one alert rather than four thousand. A heartbeat listener maintains liveness, snapshotting when a connector's gap crosses a configured threshold, which is how you distinguish a connector that is dead from one that is slow. A cron listener consumes job control messages and self-schedules recurring work, rehydrating its schedule from the database on startup.


Figure 4. The fan-out. Connectors publish and do not wait. Five consumers each own one concern, and the transfer path never passes through this plane.

Four consequences follow from that shape, and they are the reason Solace built it this way.

The monitoring plane can be entirely down while file transfers continue, because the connectors publish and do not wait for anyone. There is no synchronous coupling to fail.

Each concern scales on its own curve. Audit writes are high volume and unhurried. Alerting is low volume and urgent. They no longer share a bottleneck.

A sixth concern can be added without touching the transfer path at all.

And the operational record is produced as a byproduct of the system running, not as a separate step that someone has to remember.

That last point is why the alert throttling matters more than it looks. In a genuine site failure, every connector fails at once. An alerting design that treats each failure independently blinds the operations team at precisely the moment the declaration clock starts running.

Who is allowed to move data

Two layers of control, and they are different layers on purpose.

At Solace Event Broker, each connector holds its own client username bound to a client profile and an access control profile. It can publish to the topics it needs and subscribe to the topics it needs, under explicit limits on spool size, connection count, flow control and message size. A misconfigured connector cannot reach another connector's data, because the event broker will not let it.

At File Mesh Manager, authentication runs through local accounts, Azure AD, or SAML, alongside each other rather than as alternatives. Access tokens are short-lived and carry the holder's role and their permitted organisations and screens. Authorisation is enforced at the individual endpoint and screen, not as a broad administrator-or-user split. Credentials and identity provider secrets are stored encrypted. Everything, connectors, servers, ports, users and statistics, is scoped by organisation, so a multi-department or multi-customer deployment is isolated structurally rather than by convention.

What this is really about

Strip out the specifics and the same three decisions appear at every layer of Solace File Mesh.

The control plane is not the data plane. The thing that decides and records is separate from the thing that moves. Solution Manager never touches a byte of customer data. Neither does the Command Center. That separation is what makes the design surface safe to give to more people than you would give a production shell.

Observability is publication, not interrogation. Nothing polls anything. Components say what is happening to them, and whoever cares listens. This is why the monitoring layer can be replaced without touching the transfer path, and why the transfer path does not slow down when the monitoring layer is having a bad day.

Configuration is code. Versioned, promoted, recorded. If you cannot say what is running in production and who put it there, you do not have a platform. You have a convention and some luck.

Figure 5. Three planes. The same separation at every layer. The thing that decides and records is never the thing that moves.


None of this is specific to file transfer, and none of it is specific to one market or one regulator. The windows are closing everywhere, and the institutions under them face the same question, which is whether what they switch to is correct. Those three properties are what we look for in a platform, and why we chose to build our practice on this one, when the requirement is that something keeps working, provably, on a day when it matters and nobody has time to investigate.

The clock will start eventually, at some institution, on some Tuesday. Almost everything that determines the outcome will have happened before it does.



Iesoft Technologies is a Solace partner. We implement and operate Solace File Mesh for customers, and build event-driven systems for regulated industries, having worked with Solace and Apache Kafka for over fifteen years across capital markets, aviation, healthcare and IoT.


Solace File Mesh, the Files Micro-Integration, File Mesh Manager and Solution Manager are products of Solace Corporation. If you are looking at replication, integration or event architecture under a regulatory deadline, we would be glad to talk.

Implementing Solace File Mesh for market infrastructure: What building disaster recovery replication for market infrastructure taught us, and how Solace File Mesh answers it.

Iesoft Technologies is a Solace partner. File Mesh Solution, Includes File Micro Integration & Solution Manager for “Design, Deploy & Govern”. Iesoft Technologies implements and operates it for customers. This is what we have learned doing so where the regulator is watching.

In 2021, the regulator of one of the world's busiest securities markets rewrote its recovery clock.

Market infrastructure institutions, the exchanges, clearing corporations and depositories that a market actually runs on, had until then been allowed several hours to recover from a disaster, and a generous share of that simply to decide that one had occurred. The revised framework cut the whole thing to well under an hour: minutes to declare, and a fixed number of minutes from that declaration to be running live from the disaster recovery site. Tolerable data loss was halved. Institutions had a single quarter to comply.

It was one regulator in one market, but the direction was not local. Operational resilience regimes in Europe, Southeast Asia and the Gulf have been compressing the same windows, for the same reason: a market that is down is a market failing at the one thing it exists to do.

Most of the commentary at the time was about the switchover. That is the visible part: the moment someone senior decides, and traffic moves to another building hundreds of kilometres away.

But the recovery window is not when you copy the data. By the time that clock starts, the data has to already be there. What the rule actually demands is that every file your risk, collateral, clearing and settlement systems depend on has been arriving at a second data centre continuously, correctly, all day, on a day when nothing was wrong and nobody was watching. The recovery window is not the replication window.

It is when you find out whether the replication was any good.

We implement that layer. As a Solace partner, we deploy and operate Solace File Mesh for institutions that live under this clock. What follows is what the clock does to an engineering problem, what is inside Solace's platform that answers it, and what we have learned putting it into production for our customers.


Part one: the brief the clock writes

Under the old framework, one clause quietly permitted an enormous simplification. Updates at the primary data centre had to be reflected at the recovery site immediately, which the same sentence then qualified as before end of day.

Before end of day is a batch window. A batch window means a scheduled job, a maintenance slot, and a comfortable assumption that if tonight's run fails you have until tomorrow morning. An entire generation of replication architecture was built on that sentence, and all of it was compliant.

A window measured in minutes, end to end, is a different problem.

The shape

Solace's Files Micro-Integration replicates files from a Source to a Sink, Solace's term for the receiving end, and the two halves never speak to each other directly. The Source sits beside the source filesystem, reads files, breaks them into parts and publishes each part as a guaranteed message onto the Solace event mesh. The Sink at the far site consumes those parts from a durable queue and writes files back out onto the destination filesystem.

Putting an event broker in the middle of a file transfer looks like an odd choice until you consider what the alternative asks you to build. A point-to-point protocol between two data centres gives you one path, and every property you care about, persistence, retry, ordering, delivery guarantees, multi-site fan-out, becomes your problem to implement and your problem to prove. Publishing onto a mesh that already provides those properties means the replication logic can be about files rather than about networks. The same stream reaches a near site and a disaster recovery site without the source knowing either exists.

It also means the two ends are decoupled in time. The Sink can be down for maintenance while the Source continues to publish. Messages wait in the queue. Nothing is lost and nothing needs coordinating.

Figure 1. The shape. The Source and Sink never speak directly. Data, command-centre events and the resume checkpoint all live in Solace Event Broker, and one publish reaches every site that listens.

The problem that actually matters

Here is the scenario that shapes everything else.

It is 14:20 on a trading day. A four gigabyte collateral file is 3.8 gigabytes of the way across. The WAN link drops.

The naive answer is to start again. Under an objective measured in hours, starting again is survivable. Under one measured in minutes it is not, and on a busy afternoon it may not even be survivable in absolute terms, because starting again means re-sending 3.8 gigabytes you have already sent, over a link that has just proven it can fail.

So the Source checkpoints. Before each part goes out it records where it has reached: the file identity, the total parts, the parts sent, and the exact byte read position. On restart it fetches that checkpoint, compares the file it finds against the file the checkpoint describes, seeks to the recorded offset and continues.

The interesting decision is where the checkpoint lives. It is a Last Value Queue, a Solace Event Broker construct that retains only the most recent message on a topic. The resume state sits in the message layer itself.

Consider what that avoids. No side database to provision, secure, back up and make highly available in both data centres. No local state file that can diverge from what the broker believes. No additional piece of infrastructure that has to be up for replication to be recoverable, and therefore no additional thing that can be down at 14:20. The broker is already the most available component in the architecture, because everything else depends on it. Putting the resume pointer inside it means the checkpoint has exactly the same availability as the transport.

Most teams treat a broker as a pipe. Solace's design treats it as a place to keep things. Fifteen years of building on Solace and Kafka have taught us how rare, and how right, that instinct is.


Figure 2. 14:20. The cumulative byte count travels on every part. The checkpoint lives in a Last Value Queue on Solace Event Broker, so it has the same availability as the transport.

Trusting nothing on the receiving side

Guaranteed messaging can redeliver. That is a feature, and it is why nothing gets lost when a consumer dies mid-processing. It also means the Sink must assume that any part may arrive twice, or that a part it has already written may show up again after a reconnect.

The obvious defence is to track part numbers. The Sink does not, because part numbers describe what the sender intended, and after a resume the sender's intentions and the receiver's disk may disagree.

Instead the Sync works in bytes. It derives the incoming part's start position from the cumulative byte count and the payload length, then compares that against how far the file on disk has actually got. Four outcomes:

Contiguous. The part begins exactly where the file ends. Write it.

Overlapping. The part begins before the current end of file but extends past it. This is what a resume looks like when the Source's checkpoint is slightly behind the Sink's disk. The buffer is trimmed to the new tail and only the genuinely new bytes are written.

Gapped. The part begins after the current end of file, meaning something in between never arrived. The part is not written, because writing it would produce a file that is the right length and wrong in the middle. That is the failure mode worth fearing: a corrupt file that looks fine.

Duplicate. The part covers ground already written. Discarded.


Figure 3. Four outcomes. Every incoming part is placed against the high-water mark of the file on disk. A gap is never written.

The overlapping case is the one that repays the effort. A resume across a WAN under load is not clean. Somebody's checkpoint is always a little stale. Handling that at byte granularity rather than part granularity is what lets a transfer resume without either losing a fragment or writing one twice.

Fidelity is not the same as bytes

A replica with the right contents and the wrong ownership is not a replica. When the recovery site has to run live operations, files have to be readable by the processes that expect to read them.

So the source's ownership and permissions travel with the data and are reapplied at the destination. Ownership is carried as user and group names rather than numeric ids, and resolved locally at the far end, because the same account can hold different numeric ids on different hosts and the numeric id is the thing that would silently be wrong.

Two payload types get their own handling. A zero-byte file is sent explicitly as an empty file rather than as a part with nothing in it, so the Sink cannot mistake it for a truncated transfer. A directory is its own message type and is created as a directory, so that empty directory structure survives replication. Neither is glamorous. Both exist because directory trees replicated without them come out subtly wrong, and subtly wrong is the expensive kind.

Four ways to replicate, and why one of them matters more now

The Files Micro-Integration handles a static file, replicated once per run. A directory, single level. A directory, recursive to arbitrary depth.

And a dynamic file, which is the one the regulation has been pushing toward. Here the Source replicates to end of file and then does not stop. It keeps watching, and as the source file grows it continues to publish. The transfer is bounded by the trading day rather than by the file's current length.

That mode exists because some files are written throughout the session rather than produced at the end of it. It has become more important since, because the same regulator has more recently moved the Recovery Point Objective from a matter of minutes to near zero, with zero data loss required for clearing corporations and depositories. A few minutes of tolerable loss can be met with scheduled replication. Near zero cannot. Continuous replication of a file that is still being written stops being a special case and starts being the default.

The details that come from having been paged

Some behaviour in the Files Micro-Integration only makes sense once you have operated it through something going wrong at an unhelpful hour. These are the details we point customers at first.

The Source's directory scan looks back a bounded number of days for modified files, four by default. That bound is there because of long weekends. Without it, a holiday during which someone touched an archive directory produces a mass re-replication on the next working morning, consuming exactly the bandwidth the morning needs.

There is a skip window that ignores files created within a configured interval of the last successful run, which is how you avoid picking up a file that is still being written by something else.

There are per-directory file ceilings, duplicate path detection across configuration entries, and wildcard whitelists and blacklists, because the scope of a replication is a thing that quietly grows.

There is bandwidth limiting, because a replication that saturates the inter-site link during market hours has solved one problem by creating a worse one.

And the scheduler refuses to start an instance that is already running. It checks, finds the process, logs, and exits rather than starting a second copy. Before it launches anything at all it opens a session to the event broker to confirm the user account is enabled, so that a permissions problem surfaces as a clear pre-flight failure rather than as a replication that appears to start and then quietly does nothing.

Two streams, from the beginning

The last thing about the micro-integration is the one that turned out to matter most.

File data goes to one topic. Everything the Source and Sink have to say about themselves goes to others: a start event when a run begins, heartbeats at a configured interval, warnings, errors, a multipart event, a completion event carrying the run's totals.

Nothing polls them. They publish.

It reads as a design convenience. In practice it is the whole architecture. The data path and the story of the data path are separate by design, which means the monitoring layer can be built, changed, and taken down without touching replication. It also meant that adding a new consumer of operational events required no change to the Source or Sink at all.

Part two: what one deployment cannot teach an industry

Everything above works. We have run it in production across primary, near-site and recovery data centres, on the feeds any market regulator would name as critical: risk, collateral, clearing and settlement.

Deployed on its own, without a management layer above it, it is also a system you have to be initiated into.

The configuration lives in files, one per instance, in a directory tree with a shape you learn by being shown. Topic strings are composed by hand and have to agree at both ends. Broker client profiles and access control profiles are created by hand, one per instance, granting each exactly the topics it needs. Scheduler entries are shell scripts. Job orchestration is crontab or Control-M depending on the site. Liveness is a process check every minute.

For one institution with a team that lives inside it, this is correct. Nothing in it is wasted, and the concreteness is a virtue when something breaks at 06:40.

But the clock is not one institution's problem. It landed on every exchange, clearing corporation and depository in that market at once, with one quarter to comply. It was later extended to other systemically important participants, on the explicit reasoning that a registrar holding tens of millions of investor accounts is infrastructure too. And then the target moved again, to near zero data loss, with a new requirement for a documented reconciliation process when operations resume at another site.

You cannot hand an industry a tarball and a set of conventions. Not because the engineering does not transfer, but because the knowledge does not.

That question, how do you give this to the hundredth institution without also giving them the team, is what Solace File Mesh Manager exists to answer.

Part three: Solace File Mesh Manager

Solution Manager

Solution Manager is where the hand-edited configuration file goes to die.

A data flow becomes a first-class object: named, versioned, and built visually from a catalogue of sources, destinations and transformations. That much is table stakes and every integration product claims it. What matters is the rigour around it.

Connections are objects in their own right, reusable across flows, and they are tested before they can be saved. A connection that has never successfully connected does not become part of a flow. Once connected, the designer browses the actual schemas, tables and columns on the far side, so field mapping happens against the system as it exists rather than against documentation describing how it existed.

The database coverage tells you what kind of estate Solace built this for. PostgreSQL, Oracle, MySQL and MariaDB, SQL Server, SAP HANA, Vertica. Nobody accumulates that list building for greenfield.

Flows move in both directions: database to event mesh, and event mesh to database. Real estates need both, usually on the same afternoon.

From designed to running

There is a gap in most integration platforms where the product stops and a runbook begins. You design something, the tool congratulates you, and then a human copies artefacts to a server.

File Mesh Manager closes it. A designed flow is packaged into a deployable connector, pushed to a Git repository, and deployed over SSH to a target server with image and version tag validation. The port it will use is registered against the organisation that owns it, so two connectors on the same host cannot collide. The process is started, and its status is visible from the same console that designed it.

Promotion between environments works the same way. A flow moves from development to test to production using the target environment's own approved connections, never the source environment's, and every promotion is recorded as a transaction.

The effect is that integration configuration behaves like code. Reviewable, diffable, revertable, attributable. Git is not a storage detail here. It is the system of record for what is deployed where and who put it there.

The fan-out

Behind File Mesh Manager's Command Center dashboard sit five listener services, and they are what the Source and Sink's event stream feeds.

A single Solace event stream from the running connectors fans out to five independent consumers, each owning exactly one concern. A state event listener writes the full file-transfer lifecycle history, routing malformed payloads to an error table rather than discarding them. A file complete listener updates per-connector statistics and triggers completion notifications. An error listener consumes error events and sends alerts throttled and de-duplicated against the run identifier it cached when the run began, so that one failing run produces one alert rather than four thousand. A heartbeat listener maintains liveness, snapshotting when a connector's gap crosses a configured threshold, which is how you distinguish a connector that is dead from one that is slow. A cron listener consumes job control messages and self-schedules recurring work, rehydrating its schedule from the database on startup.


Figure 4. The fan-out. Connectors publish and do not wait. Five consumers each own one concern, and the transfer path never passes through this plane.

Four consequences follow from that shape, and they are the reason Solace built it this way.

The monitoring plane can be entirely down while file transfers continue, because the connectors publish and do not wait for anyone. There is no synchronous coupling to fail.

Each concern scales on its own curve. Audit writes are high volume and unhurried. Alerting is low volume and urgent. They no longer share a bottleneck.

A sixth concern can be added without touching the transfer path at all.

And the operational record is produced as a byproduct of the system running, not as a separate step that someone has to remember.

That last point is why the alert throttling matters more than it looks. In a genuine site failure, every connector fails at once. An alerting design that treats each failure independently blinds the operations team at precisely the moment the declaration clock starts running.

Who is allowed to move data

Two layers of control, and they are different layers on purpose.

At Solace Event Broker, each connector holds its own client username bound to a client profile and an access control profile. It can publish to the topics it needs and subscribe to the topics it needs, under explicit limits on spool size, connection count, flow control and message size. A misconfigured connector cannot reach another connector's data, because the event broker will not let it.

At File Mesh Manager, authentication runs through local accounts, Azure AD, or SAML, alongside each other rather than as alternatives. Access tokens are short-lived and carry the holder's role and their permitted organisations and screens. Authorisation is enforced at the individual endpoint and screen, not as a broad administrator-or-user split. Credentials and identity provider secrets are stored encrypted. Everything, connectors, servers, ports, users and statistics, is scoped by organisation, so a multi-department or multi-customer deployment is isolated structurally rather than by convention.

What this is really about

Strip out the specifics and the same three decisions appear at every layer of Solace File Mesh.

The control plane is not the data plane. The thing that decides and records is separate from the thing that moves. Solution Manager never touches a byte of customer data. Neither does the Command Center. That separation is what makes the design surface safe to give to more people than you would give a production shell.

Observability is publication, not interrogation. Nothing polls anything. Components say what is happening to them, and whoever cares listens. This is why the monitoring layer can be replaced without touching the transfer path, and why the transfer path does not slow down when the monitoring layer is having a bad day.

Configuration is code. Versioned, promoted, recorded. If you cannot say what is running in production and who put it there, you do not have a platform. You have a convention and some luck.

Figure 5. Three planes. The same separation at every layer. The thing that decides and records is never the thing that moves.


None of this is specific to file transfer, and none of it is specific to one market or one regulator. The windows are closing everywhere, and the institutions under them face the same question, which is whether what they switch to is correct. Those three properties are what we look for in a platform, and why we chose to build our practice on this one, when the requirement is that something keeps working, provably, on a day when it matters and nobody has time to investigate.

The clock will start eventually, at some institution, on some Tuesday. Almost everything that determines the outcome will have happened before it does.



Iesoft Technologies is a Solace partner. We implement and operate Solace File Mesh for customers, and build event-driven systems for regulated industries, having worked with Solace and Apache Kafka for over fifteen years across capital markets, aviation, healthcare and IoT.


Solace File Mesh, the Files Micro-Integration, File Mesh Manager and Solution Manager are products of Solace Corporation. If you are looking at replication, integration or event architecture under a regulatory deadline, we would be glad to talk.

Want to build something great?

Let's build something extraordinary together

Request a free consultation

Want to build something great?

Let's build something extraordinary together

Request a free consultation