Safety System Products / Open Source
Hardening a Cross-Platform BLE Stack
Universal BLE was already part of my Bluetooth work on embedded devices and their mobile and desktop applications. Then a Windows application started crashing without a Dart exception. I traced the failure into the native plugin and contributed the first fixes upstream. Manual testing led me to build a hardware-in-the-loop fixture, drawing on my functional-safety experience. With an agentic loop coordinating firmware flashing and test runs, I could vary failure timing, verify the fixes, and find more bugs in the Windows implementation.

Project Brief
- Context
- Cross-platform Flutter application for a functional-safety device
- My role
- Cross-platform BLE, root-cause analysis, Windows and Web hardening, HIL architecture and verification direction
- Platforms
- Flutter, Windows C++/WinRT, Web Bluetooth, Zephyr, nRF52
- Upstream
- Merged contributions to Universal BLE
1.Working with Bluetooth on both sides
01 It started with the Smart Insoles
I was already familiar with Bluetooth from the embedded side. The Smart Insole Firmware project also brought me into Flutter application development: two nRF52840 devices in a primary-secondary setup, live sensor streaming, stored-session synchronization, device health, configuration, authentication, and firmware updates with rollback.
The companion app had to discover and identify devices, manage connections, understand the GATT model, transfer binary data, and recover from interruptions. It also had to present two physical insoles as one product. I found Universal BLE useful here because it gave me a common Flutter API over each platform's native Bluetooth implementation.
02 Using Universal BLE at SSP
I later used the same package in an upcoming SSP application for an industrial device, working across Windows, Android, and Web. Discovery, GATT reads and writes, notifications, and reconnect behavior were already familiar from working on the firmware and application sides of BLE.
Universal BLE had saved me work across these projects. When I found a critical bug I could reproduce and investigate, I wanted to give something back to the library I was using. Contributing the fix upstream would also help other applications running into the same problem.
03 A replacement application with a higher reliability bar
The SSP application will replace a legacy application that had accumulated too many problems. It communicates over BLE with a device that is part of a functional-safety system.
The application is not a certified safety component, but it interfaces with a functional-safety device. Weak links, disconnects, and recovery therefore had to be part of normal operation.
2.A Windows crash without a Dart exception
01 Windows was the odd platform out
In the original weak-link tests, Android and Web handled the failures, while the Windows build could simply terminate. There was no Dart exception, no stack trace, and no useful failure event despite the exception handling already in the application. The process was dying below Flutter in native Windows code.
At first it happened intermittently. Then I started walking the BLE device away from the laptop. A weaker signal and repeated reconnection attempts made the crash reliably reproducible, so I could follow the failure into the native code.
02 Following the failure down the stack
On Windows, calls to the Universal BLE Dart API eventually reach a native C++ plugin built on WinRT. I traced the crash into the plugin and found WinRT exceptions escaping from asynchronous connection operations and event callbacks. Flutter never had a chance to catch them.
3.The first fixes and manual testing
The first fix
My first upstream contribution, PR #217, added exception handling to the Windows connection and notification callbacks. If WinRT now throws there, the plugin cleans up the affected connection state and reports the failure through the existing Dart connection callback instead of taking down the process.
I manually reproduced the failure and checked that the application could reconnect afterward, then submitted the fix upstream. It used the existing connection callback, so the public Dart API did not need to change.
Why Dart exception handling was not enough
A try and catch in Dart only helps if the failure crosses the platform boundary. These exceptions were thrown inside native asynchronous operations and event callbacks. The Windows process was already gone before Flutter could react.
Once the process stayed alive, I could investigate what a badly timed disconnect or reconnect did to the connection state.
4.Hardening the connection lifecycle
One connection, several asynchronous owners
As I continued testing the Windows implementation, I found that Flutter method handlers, WinRT callbacks, and asynchronous GATT operations could all touch the same connection state during disconnect or cleanup. An operation started by an old connection could even finish after a new connection had replaced it.
Native objects had to stay alive for as long as an operation used them. Work belonging to an old connection also had to lose the right to change the new one, even if it finished successfully.
Hardening ownership and stale operations
My next contribution, PR #278, added synchronized ownership of the connected-device state and a generation number for connection attempts. Every new attempt advances that number. Callbacks and asynchronous completions can then check whether they still belong to the current connection before changing anything.
I manually checked scanning, connection, service discovery, reads and writes, notification changes, and disconnects against a physical BLE device. The ownership and cleanup changes below also went through the HIL fixture once it was available. WinRT completion handling needed care too: an operation could finish successfully after its connection had already been replaced.
- Shared ownership for native connection state used by in-flight work
- Connection generations that invalidate superseded attempts and callbacks
- Synchronized installation, lookup, replacement, and removal of devices
- Cleanup ordered around handlers, subscriptions, services, and native devices
- Notification changes committed only after the native operation succeeds
5.Verifying the fixes with hardware
01 Moving beyond manual reproduction
The manual tests passed, but reproducing a failure by walking a device away from a laptop gave me little control over its timing. In functional-safety work, I was used to inserting faults deliberately and checking the response. Here I could do that by controlling the peripheral side of the BLE connection.
The maintainer had also told me that Windows testing was difficult. A physical fixture would let us repeat failures through the application and native Bluetooth stack, including the radio link.
02 Applying my functional-safety experience
I wanted to check that the fixes addressed the failures I had found, including their timing and recovery behavior. The peripheral could deliberately delay a response, reject a request, disconnect during an operation, or restart while the host still had work pending.
A passing test needed to tell me whether the fix addressed that particular failure and whether the library recovered afterward. I also wanted to run the same scenario with the protection removed. That would expose a test which passed without ever exercising the fix.
03 Building the peripheral the test needed
I designed a hardware-in-the-loop (HIL) setup around an nRF52 DK running purpose-built Zephyr firmware. The Windows test application talks to the board over BLE, while USB serial carries the firmware logs. Each test can arm a fault before starting the BLE operation it wants to disturb.
With a programmable peripheral, I could reproduce the same fault repeatedly and vary its timing. That gave me a way to test connection races that were difficult to hit consistently by hand.
6.Running the verification through an agentic loop
Connecting the application, library, and fixture
Building the fixture also meant repeatedly changing firmware, flashing the board, and running host-side tests. I realized agents could coordinate that work. They could implement a scenario, build and flash the Zephyr firmware, run the Flutter test, inspect the failure, and repeat after a correction. I could investigate more edge cases without manually carrying every change through that cycle.
I gave the agents context from three repositories: the SSP application, Universal BLE, and the Zephyr fixture firmware. This is usually called context engineering. In this case, it let a host-side test be checked against the firmware command it used, the native code it exercised, and the recovery behavior the application needed.
Different agents for different parts
I split the agent work by concern. One agent could stay inside Zephyr and the fault protocol, another could work on the Flutter integration tests, while another reviewed C++ ownership, WinRT completion, and teardown. Separate verification work then checked what came back.
Independent tasks could run in parallel, with their findings returning to one coordinating thread before the next decision. I defined the architecture, decided what could be delegated safely, connected the results, and rejected changes that did not match the physical system.
Checking that a test exercises the fix
I organized that work into an agentic loop, sometimes called loop engineering. Each cycle started with a concrete failure scenario. The agents worked on the matching host test and peripheral behavior, built and flashed the fixture, ran the test, and used the result to correct the implementation or the test itself.
The same loop checked the hardening. I would remove or bypass a fix and have the scenario run again. If it stayed green, I needed to find out why it had missed the missing protection. Once it failed for the expected reason, I restored or improved the fix and checked that the test passed, including a follow-up read, write, reconnect, or resubscription.
My implementation and review responsibilities
I used AI heavily here. Flutter integration testing was new to me, so the agents handled much of the detailed host-side implementation. They also produced the first version of the Zephyr fixture firmware and automated parts of the build and flash process. I reviewed, corrected, and improved the firmware work because embedded software, BLE, and Zephyr are the areas where I already knew what good looked like.
I defined the fault model, the system boundaries, what counted as a correct result, how recovery should behave, and how the three repositories fit together. I also checked whether an implemented fault represented the failure class under investigation and whether its passing test was meaningful.
Extending the loop to native Windows tests
The same approach later supported native Windows tests and startup-close reproduction, alongside the HIL and fault-injection runs. That mattered when shutdown testing reached a failure during Bluetooth initialization.
7.A physical fault-injection system for BLE
Testing from Flutter to the physical peripheral
The HIL system merged in PR #282 starts with a Flutter integration test calling the Universal BLE Dart API. Like a request from the SSP application, it passes through the operation queue, Pigeon channel, Windows C++ plugin, WinRT, and Windows Bluetooth stack. The request then travels over the radio link to the Zephyr GATT server on the nRF52 DK.
The physical peripheral lets the tests exercise operating-system Bluetooth behavior, native asynchronous completions, and ATT responses over the radio link. Those parts of the system were missing from software-only tests.

Baseline behavior before fault injection
The original suite had 64 tests: 25 for baseline BLE behavior and 39 for fault injection. The baseline tests cover scanning and advertisement data, service and characteristic discovery, reads and writes with exact payload verification, notifications and indications, subscription changes, MTU reporting, disconnects, reconnects, repeated connection cycles, and concurrent operations.
Before injecting faults, I needed to know that ordinary BLE operations worked. These tests also provide a clean read, write, or connection operation to run after a failure and check whether the library actually recovered.
Injecting faults from the peripheral
The fault-injection tests arm the Zephyr peripheral before starting an otherwise normal BLE operation. The board can return ATT errors, delay a response past the timeout, disconnect with a request pending, complete stale work after a reconnect, change its GATT database, or send notifications with gaps, duplicates, reordered sequence numbers, and mixed payload sizes.
Most scenarios then perform another read, write, reconnect, or subscription. A library can report the injected error correctly and still leave stale state that breaks the next operation. The follow-up operation checks for that.
8.Finding more Windows bugs
GATT objects and the first notification
Running faults repeatedly through the fixture exposed more Windows problems. Those fixes became PR #284. In-flight reads, writes, discovery, and descriptor operations needed to retain their native GATT objects. Service changes also needed to refresh the GATT map safely and restore handlers for existing subscriptions.
One example was the first notification: the peripheral could send its first notification while the Client Characteristic Configuration (CCC) descriptor write was still completing. Windows registered the handler too late and lost that value. I moved registration before the enable write, with rollback if the write failed. The same PR corrected descriptor return values and copied notification payloads when their underlying buffer also contained platform-message bytes.
Queued callbacks and plugin shutdown
Connection events also needed their generation checked when delivered on the UI thread. Checking only when an event was queued still allowed an old disconnect event to arrive after a newer connection had started. I added a reconnect test that swept short delays around a pending disconnect callback.
For shutdown, a callback tracker stopped new work from entering and kept plugin state alive until callbacks already running had finished. The drain dispatched COM continuations and Windows messages so work waiting for the UI thread could finish. The baseline and fault-injection suites passed with that drain in place. Later, closing the application during startup exposed a different shutdown failure.
9.Connecting immediately after discovery
Immediate connection after scanning exposed another Windows failure, addressed in PR #311: a device could be available immediately after scanning, while the first uncached GATT service query returned Unreachable. Creating the native device object did not itself establish the BLE connection. The query was starting that connection and could reach it before it was ready.
I added two retries for that specific status, after 250 and 500 ms. Authorization and protocol errors still fail immediately, and cached services are never used as proof of a live connection. Native failure messages also map to specific Dart error codes, with exact matching so unrelated errors are not misclassified.
The added HIL scenario repeatedly scans, connects immediately, discovers services, and reads a characteristic. It bypasses the fixture helper's reconnect retries so a failed first attempt cannot be hidden by the test itself.
10.Reconnect problems on Web
01 Advertisement watching overlapped with connection setup
Rapid connection cycles and auto-reconnect exposed problems on Web too. In PR #310, I made GATT connection setup wait for advertisement watching to stop. Pending watcher startup is aborted before cleanup waits for it, and startup and stop operations each have a five-second bound. A successful watch can still continue beyond those five seconds.
Connection generations reject old events and discovery results before they can repopulate the service cache. Teardown cancels connection and characteristic subscriptions. The disconnected event is published after the disconnect reservation is released, so an application callback can start a reconnect.
02 Cancellation had to leave room for a retry
The review caught an awkward case: a browser connection operation might never settle after cancellation. Keeping its reservation forever would block every later attempt. I added a five-second cancellation grace period, and made late completions check whether a newer pending attempt or tracked connection owns the link before disconnecting it.
The dependency wrapper did not expose the caller-owned abort signal needed to cancel advertisement watching. I used its native advertisement API for that operation. GATT connection and discovery continue through the wrapper.
The PR also preserves typed BLE error codes through exception wrapping, with a regression test.
11.Closing Windows while Bluetooth was still starting
Radio initialization during shutdown
PR #319 addressed a crash when the app closed while Radio::GetRadiosAsync() was pending. Initialization held a lease on the general callback tracker. Shutdown pumped UI messages while waiting for its continuation, and those messages could enter Flutter after engine teardown had begun.
For this fix, Codex (GPT 6.1 Sol) handled implementation and test execution entirely. Initialization was separated from that tracker and given a WinRT completion handler. Pending radio enumeration no longer holds a lease. Only its short completion and UI publication are tracked, late work is ignored after shutdown starts, and an already-entered initialization completion drains without pumping application messages. The cross-thread notification window handle is atomic too.
Reproducing the close in the real runner
Before the fix, every startup-close reproduction run crashed with 0xC000041D. Afterward, repeated runs in both Debug and Release exited cleanly when closed during the first two seconds of startup.
Verification covered native tests, startup-close reproduction, and the Windows HIL and fault-injection runs. A native test checks the initialization lifetime protocol with a controlled WinRT operation. Using the general callback tracker as a countercheck fails its message-preservation assertion. A separate script exercises startup and close in the actual Flutter runner.
12.What changed upstream
01 A test cycle I could keep using
I started with a crash I could reproduce by weakening the radio link. By the later fixes, I could specify the failure timing in firmware, run it repeatedly, and check the next BLE operation for recovery. The agents made that cycle quick enough to use while investigating new bugs. Checking each test with its protection removed also helped me catch gaps in the verification itself.
02 A hardware fixture for future Windows changes
The maintainers plan to run the hardware scenarios before releases. The fixture gives them a way to reproduce timing failures and check cleanup, reconnection, and subsequent BLE operations against a physical peripheral.
03 Contributing back to Universal BLE
By early October 2026, my account, usmanmehmood55, ranked second in Universal BLE's contributor graph for the preceding few months. What started as a fix for a crash in my application had grown into ongoing work on the library and its hardware verification tools.

04 Upstream contributions
PR #217 contained the native exceptions and returned failures to Flutter. PR #278 fixed ownership, asynchronous lifetime, reconnect, and notification-state problems. PR #282 added the HIL and fault-injection system used to reproduce and verify them.
PR #284 extended the Windows GATT and shutdown protection. PR #311 addressed transient discovery failures, followed by PR #310 for Web connection lifecycle problems and PR #319 for the startup-close crash. All seven are merged.
13.Test scope and current limits
The automated hardware suite targets Windows and requires an nRF52 DK with exclusive access to a Bluetooth adapter. It does not run in ordinary pull-request CI yet. I am working on automating these tests on the physical hardware.
The Zephyr fixture injects GATT and ATT errors, timing faults, disconnects, service changes, and notification anomalies. It is a standards-compliant peripheral and cannot generate arbitrary malformed link-layer packets.
The Web lifecycle changes were manually tested; automated browser-race coverage is still missing. The Windows startup-close fix covers initialization. Shutdown during other active BLE operations still uses the general message-pumping drain.