Debugging a Bluetooth Kernel Regression on NixOS

After a routine nix flake update on my ASUS ZenBook S13 (NixOS, MT7922 combo WiFi+BT chip), Bluetooth stopped working entirely. No devices visible, no controller available. dmesg gave one line:

Bluetooth: hci0: Failed to send wmt func ctrl (-22)

What followed was two days of chasing wrong hypotheses before finding the real culprit — a two-line bounds check added three days before the kernel version my system had just upgraded to.


The Setup

  • Machine: ASUS ZenBook S13 running NixOS
  • Chip: MediaTek MT7922 — a combo chip where WiFi (mt7921e, PCIe) and Bluetooth (btusb/btmtk, USB) share a WMT (Wireless Module Technology) bus
  • Kernel before update: 6.18.21
  • Kernel after update: 6.18.32

Wrong Lead #1 — USB Autosuspend

The error message wmt func ctrl (-22) is -EINVAL in USB terms, and the WMT bus is a USB-adjacent protocol. My first search led to Red Hat Bugzilla #2401216 and a known regression from commit 5c5e8c52e3ca ("Bluetooth: btmtk: move btusb_mtk_[setup, shutdown] to btmtk.c") which had removed usb_autopm_get_interface() guards from the btmtk setup and shutdown routines.

This was fixed by commit 67dba2c28fe0 in December 2024. Checking the 6.18.32 source confirmed the fix was already present — so this wasn't the cause.

Still, I applied the common workaround just in case:

boot.extraModprobeConfig = "options btusb enable_autosuspend=n";

$ cat /sys/module/btusb/parameters/enable_autosuspend
N

The parameter was active. Bluetooth was still broken.


Wrong Lead #2 — WMT Bus Ordering Race

The MT7922 is a combo chip: WiFi claims the WMT bus via PCIe (mt7921e), Bluetooth via USB (btusb/btmtk). The theory: if mt7921e initialises first and grabs the bus, btmtk can't send the wmt func ctrl command.

The fix is to load btusb from the initrd so it wins the ordering race:

boot.initrd.kernelModules = [ "btusb" ];

Looking at the boot log more carefully though:

Bluetooth: hci0: Failed to send wmt func ctrl (-22)   ← 02:51:47
...
mt7921e 0000:01:00.0: ASIC revision: 79220010         ← 02:51:55

mt7921e doesn't even appear until 8 seconds after the Bluetooth failure. This was a separate real issue (worth fixing), but not what broke things here.


Finding the Real Culprit

With both theories ruled out, I went back to basics: what actually changed in the kernel?

The nixpkgs update bumped the kernel from 6.18.21 → 6.18.32. (The nixpkgs-old pin I use for linux-firmware has a much older kernel — 6.12.40 — which also predates the regression and thus works fine.)

I checked btmtk.c at both versions via the GitHub API against gregkh/linux:

gh api "repos/gregkh/linux/commits?path=drivers/bluetooth/btmtk.c&sha=v6.18.32&per_page=10"

The relevant entry in the log:

2026-04-21  624fb79dadc1  Bluetooth: btmtk: validate WMT event SKB length before struct access

This commit landed between 6.18.21 and 6.18.32. Reading the diff confirmed it as the culprit.

What the Commit Changed

In btmtk_usb_hci_wmt_sync(), the BTMTK_WMT_FUNC_CTRL response handler was changed from:

// 6.18.21 — direct cast, no bounds check
case BTMTK_WMT_FUNC_CTRL:
    wmt_evt_funcc = (struct btmtk_hci_wmt_evt_funcc *)wmt_evt;
    if (be16_to_cpu(wmt_evt_funcc->status) == 0x404)
        status = BTMTK_WMT_ON_DONE;
    else if (be16_to_cpu(wmt_evt_funcc->status) == 0x420)
        status = BTMTK_WMT_ON_PROGRESS;
    else
        status = BTMTK_WMT_ON_UNDONE;
    break;

to:

// 6.18.32 — bounds check added
case BTMTK_WMT_FUNC_CTRL:
    if (!skb_pull_data(data->evt_skb,
                       sizeof(wmt_evt_funcc->status))) {
        err = -EINVAL;        // ← this is what breaks MT7922
        goto err_free_skb;
    }
    wmt_evt_funcc = (struct btmtk_hci_wmt_evt_funcc *)wmt_evt;
    if (be16_to_cpu(wmt_evt_funcc->status) == 0x404)
        status = BTMTK_WMT_ON_DONE;
    ...

The skb_pull_data call tries to pull the 2-byte status field from the socket buffer. On MT7922 hardware, the firmware omits this field entirely — it sends a shorter response than the driver now expects. The bounds check correctly detects the short packet, but then returns a hard error instead of handling it gracefully.

Why the Old Code "Worked"

In 6.18.21 the code cast wmt_evt directly to wmt_evt_funcc without bounds checking — reading potentially garbage bytes beyond the end of the actual response. Since garbage is unlikely to equal 0x404 (ON_DONE) or 0x420 (ON_PROGRESS), the driver fell through to BTMTK_WMT_ON_UNDONE and returned 0 (success). The Bluetooth stack continued, and the chip worked fine. Technically undefined behaviour, practically harmless.


Confirming the Root Cause

To be sure, I temporarily downgraded to the old kernel via my existing nixpkgs-old pin:

boot.kernelPackages = inputs.nixpkgs-old.legacyPackages.${pkgs.system}.linuxPackages;

This gave me 6.12.40 — predating 624fb79dadc1 entirely. Bluetooth worked immediately.


The Fix

The fix is two lines: instead of returning -EINVAL when the status field is absent, treat it as BTMTK_WMT_ON_UNDONE (the same outcome the old code produced) and continue.

 case BTMTK_WMT_FUNC_CTRL:
     if (!skb_pull_data(data->evt_skb,
                        sizeof(wmt_evt_funcc->status))) {
-        err = -EINVAL;
-        goto err_free_skb;
+        status = BTMTK_WMT_ON_UNDONE;
+        break;
     }

In NixOS this is applied cleanly via boot.kernelPatches in the host config:

boot.kernelPatches = [
  {
    name = "btmtk-fix-func-ctrl-short-response";
    patch = ./patches/btmtk-fix-func-ctrl-short-response.patch;
  }
];

After rebuilding the kernel (about 45 minutes) and rebooting, Bluetooth was fully functional on 6.18.32 with the patch applied.


The Upstream Fix — Already There

Before submitting the patch, I checked the bluetooth-next tree (the upstream staging tree where Bluetooth patches land before mainline):

git clone https://git.kernel.org/pub/scm/linux/kernel/git/bluetooth/bluetooth-next.git
git log --oneline drivers/bluetooth/btmtk.c | head -5

162b1adeb057  Bluetooth: btmtk: accept too short WMT FUNC_CTRL events
041e88fb0c08  Bluetooth: btmtk: validate WMT event SKB length before struct access

Pauli Virtanen (pav@iki.fi) had already found and fixed it — three days after the regression landed, on April 24, 2026. His commit message:

MT7925 (USB ID 0e8d:e025) on fw version 20260106153314 sends WMT FUNC_CTRL events that are missing the status field. [...] Fix the regression by interpreting too short packet as status BTMTK_WMT_ON_UNDONE, which makes the device work normally again.

Identical diagnosis, identical fix. He hit it on an MT7925; I hit it on an MT7922. Mikhail Gavrilov provided a Tested-by on MT7922 (0489:e0e2) — my exact USB ID.

Pauli likely caught it so quickly because he tracks bluetooth-next or linux-next directly on a rolling distro, so he saw the regression commit before it reached any stable release. I caught it later because NixOS pulled it in via the 6.18 stable backport to 6.18.32.

The fix is in Linux 7.1-rc1 and was backported to 6.18.33 (released 2026-05-23), the day before I finished debugging. It will arrive automatically on the next nix flake update once nixpkgs bumps its kernel version. At that point the local boot.kernelPatches workaround can be dropped.


Summary

DateEvent
2026-04-21041e88fb0c08 merges into linux-6.18.y — bounds check added
2026-04-21624fb79dadc1 — same commit in nixpkgs kernel source
2026-04-24Pauli Virtanen submits fix 162b1adeb057 to bluetooth-next
2026-04-24Fix backported to linux-6.18.y as 527cb4a55155
2026-05-23Linux 6.18.33 released with the fix
2026-05-23nixpkgs updated to 6.18.32 — regression lands on my system
2026-05-25Root cause identified, local patch applied and confirmed

Lessons:

  • When debugging kernel regressions, git log --oneline <file> between two kernel versions is one of the most direct tools available.
  • A "safety fix" (bounds checking) can introduce a regression if the firmware doesn't comply with the newly enforced constraint.
  • On NixOS, boot.kernelPatches is a practical way to carry a local kernel fix until upstream stabilises.
  • By the time you've debugged something obscure at this level, there's a real chance someone upstream already found and fixed it — check the subsystem tree before submitting.