| Age | Commit message (Collapse) | Author |
|
tbt_altmode_remove() drops the plug and cable references without
draining tbt->work. The work function dereferences those references,
and can also requeue itself in its error path. The VDM callbacks can
queue the same work item.
Disable and drain tbt->work before dropping the references. This waits
for an existing invocation and prevents subsequent schedule_work()
calls from queueing it during teardown.
This issue was found by an in-house static analysis tool and confirmed
by manual code review.
Fixes: 100e25738659 ("usb: typec: Add driver for Thunderbolt 3 Alternate Mode")
Cc: stable@vger.kernel.org
Assisted-by: Codex:gpt-5.6
Signed-off-by: Fan Wu <fanwu01@zju.edu.cn>
Acked-by: Heikki Krogerus <heikki.krogerus@linux.intel.com>
Link: https://patch.msgid.link/20260802014959.416687-1-fanwu01@zju.edu.cn
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
Since removal of the legacy tunnel port types, there are no more
users for these functions outside the main openvswitch module.
Functions to register vport_ops are also not exported. Allocating
vports without operations doesn't make a lot of sense.
Highlighted by Sashiko as a follow up to the removal of the module
infrastructure.
Signed-off-by: Ilya Maximets <i.maximets@ovn.org>
Reviewed-by: Aaron Conole <aconole@redhat.com>
Link: https://patch.msgid.link/20260812122007.457136-1-i.maximets@ovn.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
The ovskey flow-string parser has no OVS_KEY_ATTR_SCTP entry, so a
flow string containing sctp(src=.../dst=...) parses without error but
silently drops the L4 key. The resulting flow carries only
ipv4(proto=132), and the kernel rejects it: match_validate() in
flow_netlink.c requires OVS_KEY_ATTR_SCTP when the IP protocol is
IPPROTO_SCTP and returns -EINVAL for the missing key.
Register OVS_KEY_ATTR_SCTP in the parse table and add a matching
selftest that verifies SCTP flow key matching (sctp src/dst port).
One listener serves the whole test. socat's fork option handles each
association in a child, so the flow rules are the only thing that
changes between the three phases and the listener is never restarted
underneath them. -t 1 bounds how long a forked child lingers after
its association closes, and the existing kill -TERM of the captured
pid on teardown removes the listener itself.
Also enable CONFIG_IP_SCTP in the selftest kernel config. The config
checker strips underscores before comparing keys, so the entry sorts
before CONFIG_IPV6 rather than after it.
Signed-off-by: Minxi Hou <houminxi@gmail.com>
Reviewed-by: Aaron Conole <aconole@redhat.com>
Link: https://patch.msgid.link/20260811181645.1918420-1-houminxi@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
xHCI 1.0 allowed these pointers to be zero. Some Intel chipsets from the
era usually set it to zero, but sometimes (apparently) to the next TRB
after the one referenced by the previous transfer event on the endpoint.
Usually that's indeed the missed TD, but it may also be the last TRB of
a two-TRB TD already completed with Short Packet on its first TRB. Then
the driver skips all pending TDs, failing to find a match.
When handling Missed Service Error, scan TD list twice and only really
skip TDs in the second pass if the first pass found a match. This won't
catch bogus pointers to wrong TDs, but such a bug would be practically
impossible to detect automatically and isn't known to exist.
Reported-by: Bart Nagel <bart@tremby.net>
Closes: https://lore.kernel.org/linux-usb/al_hchyOdPoPWKEo@spiral/
Suggested-by: Mathias Nyman <mathias.nyman@linux.intel.com>
Fixes: d0b619599e52 ("usb: xhci: Expedite skipping missed isoch TDs on modern HCs")
Cc: stable@vger.kernel.org
Signed-off-by: Michal Pecio <michal.pecio@gmail.com>
Signed-off-by: Mathias Nyman <mathias.nyman@linux.intel.com>
Link: https://patch.msgid.link/20260806142113.2436238-18-mathias.nyman@linux.intel.com
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
handle_port_status() drops every USB3 port event when xhci->shared_hcd is
NULL. The check dates from a time when xhci-plat always created a shared
hcd, so a NULL one could only mean the hcd had been removed.
Since commit 4736ebd7fcaf ("usb: host: xhci-plat: omit shared hcd if
either root hub has no ports") that is no longer true. A controller whose
USB2 root hub has no ports gets a single roothub, the USB3 rhub is served
by the main hcd, and shared_hcd stays NULL for the lifetime of the device.
Every SuperSpeed port event is then thrown away as bogus behind a debug
message, so devices never enumerate even though the port sees the device
and its change bits stay set:
0x006a1203 Powered Connected Enabled Link:U0 PortSpeed:4
Change: CSC WRC PRC PLC
Broadcom Northstar is such a controller. USB3 works there up to 5.15 and
stops working from 5.19 onwards.
Ask xhci_get_usb3_hcd() instead. It returns the shared hcd when there is
one, the main hcd when the USB2 root hub has no ports, and NULL once the
shared hcd is gone, which keeps the original meaning of the check.
Tested on an Asus RT-N18U (BCM47081), which has a single roothub. Before
the change nothing enumerates on the USB3 port; after it SuperSpeed
devices enumerate normally over repeated connect and disconnect cycles,
the change bits shown above clear, and USB2 is unaffected on both ports.
Fixes: 4736ebd7fcaf ("usb: host: xhci-plat: omit shared hcd if either root hub has no ports")
Cc: stable@vger.kernel.org
Signed-off-by: Semih Baskan <strst.gs@gmail.com>
Signed-off-by: Mathias Nyman <mathias.nyman@linux.intel.com>
Link: https://patch.msgid.link/20260806142113.2436238-17-mathias.nyman@linux.intel.com
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
Non-ASCII characters trigger git send-email to prompt for encoding on each
modification near them, which is unnecessary and annoying.
Using plain ASCII avoids these prompts and does not change its meaning.
This change only affects comments and has no functional impact.
Signed-off-by: Niklas Neronin <niklas.neronin@linux.intel.com>
Signed-off-by: Mathias Nyman <mathias.nyman@linux.intel.com>
Link: https://patch.msgid.link/20260806142113.2436238-16-mathias.nyman@linux.intel.com
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
The 'xhci_virt_ep' struct currently contains a pointer to its parent
'xhci_hcd' struct. Since all endpoint-related structs are contained
within 'xhci_hcd', this pointer is redundant.
Remove the 'xhci' pointer from 'xhci_virt_ep' and instead pass it
explicitly to functions that require it, as some already do it.
This change reduces unnecessary complexity and aligns the code with
the rest of the xhci driver.
Memory impact:
For each device connected a struct 'xhci_virt_device' is allocated,
this struct conatains a 31 slot array of struct 'xhci_virt_ep'.
A USB hub consumes 1 slot, but every downstream device consumes
another slot.
This means that the total memory saved buy this patch is:
Devices * 31 * 8 bytes
Signed-off-by: Niklas Neronin <niklas.neronin@linux.intel.com>
Signed-off-by: Mathias Nyman <mathias.nyman@linux.intel.com>
Link: https://patch.msgid.link/20260806142113.2436238-15-mathias.nyman@linux.intel.com
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
The function ring_doorbell_for_active_rings() rings the doorbell
for any rings with pending URBs. It has a trivial wrapper,
xhci_ring_doorbell_for_active_rings(), which takes the same
arguments and simply calls the former.
Since the wrapper adds no functionality, remove it and rename
ring_doorbell_for_active_rings() to xhci_ring_doorbell_for_active_rings().
Signed-off-by: Niklas Neronin <niklas.neronin@linux.intel.com>
Signed-off-by: Mathias Nyman <mathias.nyman@linux.intel.com>
Link: https://patch.msgid.link/20260806142113.2436238-14-mathias.nyman@linux.intel.com
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
Simplify by replace BIT(0) call with its relevant macro.
Signed-off-by: Niklas Neronin <niklas.neronin@linux.intel.com>
Signed-off-by: Mathias Nyman <mathias.nyman@linux.intel.com>
Link: https://patch.msgid.link/20260806142113.2436238-13-mathias.nyman@linux.intel.com
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
This patch aims to unify the format of register macros and masks within
the xHCI driver. Currently, register macros have inconsistent bit-field
masks, get macros, and set macros, with varying naming conventions and
functionalities.
==================== Proposal ====================
* Introduce a standardized approach by using only mask macros for each bit
field, leveraging GENMASK() for enhanced clarity.
#define HCC_MAX_PSA GENMASK(15, 12)
* Utilize FIELD_GET() and FIELD_PREP() macros directly in the C code for
getting and setting values, ensuring consistency and readability.
u32 psa = FIELD_GET(HCC_MAX_PSA, reg);
* Maintain exceptions for macros that perform custom operations.
#define CTX_SIZE(_hcc) (_hcc & HCC_64BYTE_CONTEXT ? 64 : 32)
* Note, while FIELD_*() macros are beneficial, I am not suggesting that
they should always be used. Instead, use them where they simplify the
code and eliminate the necessity for custom get/set macros.
In the example below, additional FIELD_PREP() or FIELD_MODIFY() is not
beneficial.
#define HCS_MAX_SCRATCHPAD(p) (FIELD_GET(HCS_MAX_SP_HI, (p)) << 5 | \
FIELD_GET(HCS_MAX_SP_LO, (p)))
==================== Improvements ====================
Simplified Macros:
By reducing custom macros, the code becomes more straightforward.
Macros FIELD_GET() and FIELD_PREP() are commonly used, which contributes
to the code readability and consistency.
$ git grep -n 'FIELD_GET' | wc -l
9027
$ git grep -n 'FIELD_PREP' | wc -l
15407
Consistent Return Type:
All bit macros will return unsigned 64-bit values, mitigating potential
cross-architecture issues.
Unified Bit Range Definition:
The mask macro will define bit ranges, eliminating separate definitions
for get/set macros. Because, FIELD_GET() & FIELD_PREP() use mask macro.
Cleaner header file with less macros:
Fewer macros result in a cleaner and more manageable header file.
Signed-off-by: Niklas Neronin <niklas.neronin@linux.intel.com>
Signed-off-by: Mathias Nyman <mathias.nyman@linux.intel.com>
Link: https://patch.msgid.link/20260806142113.2436238-12-mathias.nyman@linux.intel.com
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
xhci_gen_setup() locates the operational registers using the capability
length read from the very first register:
xhci->op_regs = hcd->regs +
HC_LENGTH(readl(&xhci->cap_regs->hc_capbase));
If the controller is dead or has dropped off the bus, that read returns
~0, HC_LENGTH() truncates it to 0xff, and op_regs ends up 0xff bytes
past the page-aligned MMIO base, i.e. unaligned. The first access
through it, xhci_halt() -> xhci_handshake() reading op_regs->status, is
then an unaligned readl() on device memory. arm64 faults on unaligned
device accesses, so instead of xhci_handshake() catching the all-ones
value and returning -ENODEV, setup oopses:
xhci-pci-renesas 0005:08:00.0: Unable to change power state from D3cold to D0, device inaccessible
xhci-pci-renesas 0005:08:00.0: xHCI Host Controller
xhci-pci-renesas 0005:08:00.0: new USB bus registered, assigned bus number 1
Unable to handle kernel paging request at virtual address ffff80030a770103
ESR = 0x0000000096000021
FSC = 0x21: alignment fault
Internal error: Oops: 0000000096000021 [#1] SMP
pc : xhci_halt [xhci_hcd]
Call trace:
xhci_halt
xhci_gen_setup
xhci_pci_setup
usb_add_hcd
usb_hcd_pci_probe
xhci_pci_common_probe
xhci_pci_renesas_probe
This was hit with a Renesas uPD720201 that failed to power up ("Unable
to change power state from D3cold to D0, device inaccessible") yet still
reached the HCD probe path.
Read the capability register once, and if it reads back the all-ones
value (as xhci_handshake() and xhci_reset() already test for), abort
setup with -ENODEV before op_regs is derived from it. Reading it once
also avoids re-reading a register that may change under a concurrent
hot-removal.
Fixes: 66d4eadd8d06 ("USB: xhci: BIOS handoff and HW initialization.")
Cc: stable@vger.kernel.org
Signed-off-by: Breno Leitao <leitao@debian.org>
Signed-off-by: Mathias Nyman <mathias.nyman@linux.intel.com>
Link: https://patch.msgid.link/20260806142113.2436238-11-mathias.nyman@linux.intel.com
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
idr_destroy() is already called on error paths in dbc_tty_init(). Do not
call it again on exit. For symmetry with the init side, also use
IS_ERR_OR_NULL() to gate the exit steps.
Signed-off-by: Lucas De Marchi <ldemarchi@nvidia.com>
Signed-off-by: Mathias Nyman <mathias.nyman@linux.intel.com>
Link: https://patch.msgid.link/20260806142113.2436238-10-mathias.nyman@linux.intel.com
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
Make sure to set dbc_tty_driver to NULL to match the check in
dbc_tty_exit(). For that, make detached error handling path common to the
other branch in the same function.
Fixes: 4521f1613940 ("xhci: dbctty: split dbc tty driver registration and unregistration functions.")
Cc: stable@vger.kernel.org # v5.10
Cc: Mathias Nyman <mathias.nyman@linux.intel.com>
Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Signed-off-by: Lucas De Marchi <ldemarchi@nvidia.com>
Signed-off-by: Mathias Nyman <mathias.nyman@linux.intel.com>
Link: https://patch.msgid.link/20260806142113.2436238-9-mathias.nyman@linux.intel.com
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
If tty_register_driver() fails, it drops the reference, but fails to set
the global dbc_tty_driver to NULL, causing the unregister to be called
again when module exits.
On module unload dbc_tty_exit() only gates its cleanup on the driver
pointer being non-NULL, so it operates on the already-freed driver:
module_init(xhci_hcd_init)
xhci_hcd_init()
xhci_dbc_init() [return value ignored]
dbc_tty_init()
tty_register_driver() fails
tty_driver_kref_put() -> driver freed
(dbc_tty_driver left dangling)
...
module_exit(xhci_hcd_fini)
xhci_hcd_fini()
xhci_dbc_exit()
dbc_tty_exit()
if (dbc_tty_driver) -> true (dangling)
tty_unregister_driver() -> use-after-free
Fixes: 4521f1613940 ("xhci: dbctty: split dbc tty driver registration and unregistration functions.")
Cc: stable@vger.kernel.org # v5.10
Cc: Mathias Nyman <mathias.nyman@linux.intel.com>
Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Signed-off-by: Lucas De Marchi <ldemarchi@nvidia.com>
Signed-off-by: Mathias Nyman <mathias.nyman@linux.intel.com>
Link: https://patch.msgid.link/20260806142113.2436238-8-mathias.nyman@linux.intel.com
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
If a ring stops on a TD that is about to be cancelled then the xHC ring
hardware dequeue pointer needs to move past the TD to flush TRBs from
xHC cache.
The TRB after the cancelled TD might be a no-op TRB, or a link TRB.
Moving the dequeue to a link TRB has caused isses on some hosts, and
moving it to a no-op TRB can be an issue for control endpoints as
xhci specification 4.8.3 'Endpoint Context State" states that
The Default Control Endpoint shall return to the Running state when the
Doorbell is rung for the next Setup Stage TD sent to the endpoint.
Solve this by always moving the dequeue pointer to the next valid
TD. If ring is empty and there are no queued TDs then move the dequeue
pointer to the enqueue pointer.
If enqueue points to a link TRB on a empty ring then propagate enqueue
to next segment before pointing dequeue to it.
Note that this patch ended up almost identical to a simplifiaction patch
done earlier by Michal Pecio, see link. That patch was not added due to a
potential, somewhat theoretical issue of moving dequeue backwards.
Turns out improving cancelled control transfers end up with the same code,
and is now worth taking.
Code is very likely subconsciously based the patch by Michal Pecio.
Link: https://lore.kernel.org/linux-usb/20250225125939.7a248e38@foxbook/
Signed-off-by: Mathias Nyman <mathias.nyman@linux.intel.com>
Link: https://patch.msgid.link/20260806142113.2436238-7-mathias.nyman@linux.intel.com
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
Avoid all extra endpoint state changes after the roothub link
is lost due to disconnect or link error, and endpoint is known
to be in a non-running state.
Rapid endpoint state changes involving endpoint reset, restart, and
stopping the endpoint have caused xHC failures to complete stop
endpoint command. xhci driver sees this as a fatal flaw and tears
down xhci.
These endpoint state changes are normally part of recovery from
transaction errors or URB cancel.
In this case recovery is not needed.
Add an endpoint state called EP_DROP_PENDING.
Set ep->ep_state |= EP_DROP_PENDING when an endpoint is found in a
halted or stopped non-running state, and the roothub link is
lost. Prevent endpoint from restarting.
URB cancel doesn't need to stop the endpoint if EP_DROP_PENDONG is set.
URBs can be given back directly.
Endpoint is, and will remain stopped until it's dropped.
Tested-by: Xu Rao <raoxu@uniontech.com>
Signed-off-by: Mathias Nyman <mathias.nyman@linux.intel.com>
Link: https://patch.msgid.link/20260806142113.2436238-6-mathias.nyman@linux.intel.com
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
Prevent transfer retry and endpoint recovery if the device or its parent
disconnected from the roothub. Just like link error case.
There is a suspicion some xHC controllers may stop processing endpoint
related commands after the last USB device disconnects from the host.
Disconnect often causes transaction errors, xhci driver tries to (soft)
reset and restart the endpoint to recover it.
Hub driver again will cancel all pending URBs once disconnect is detected,
stopping the endpoint right after (soft) reset restarted it.
xHC controller sometimes fail to complete the stop endpoint command,
leading to driver timing out, and tearing down xhci
Prevent extra endpoint (soft) reset after xhci driver is aware of the
parent roothub port disconnect.
Tested-by: Xu Rao <raoxu@uniontech.com>
Signed-off-by: Mathias Nyman <mathias.nyman@linux.intel.com>
Link: https://patch.msgid.link/20260806142113.2436238-5-mathias.nyman@linux.intel.com
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
Driver already prevents useless transfer retry and endpoint recovery
for devices directly connected to a root port with link errors.
These devices are either disconnecting or will be reset. Link is gone.
Move the flag indicating link error from the xhci device structure to
the root port strucure, allowing all child devices behind hubs to easily
check for root port link errors, avoiding useless transfer retries and
endpoint recovery.
This extends the previous endpoint recovery prevention in
commit b8c3b718087b ("usb: xhci: Don't try to recover an endpoint if port
is in error state.")
Only root port link errors can be detected early by xhci driver,
not link errors between external hubs and their children.
Tested-by: Xu Rao <raoxu@uniontech.com>
Signed-off-by: Mathias Nyman <mathias.nyman@linux.intel.com>
Link: https://patch.msgid.link/20260806142113.2436238-4-mathias.nyman@linux.intel.com
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
The frame id field can be set for the first TD of the first isoc
URB to schedule the start of an isoc stream even in host doesn't
support CFC (Contiguous Frame ID Capability)
Set the frame ID TRB field of the first isoc TD unless URB has the
schedule immediately 'URB_ISO_ASAP' transfer flag set.
cc: Dylan Robinson <dylan_robinson@motu.com>
Signed-off-by: Mathias Nyman <mathias.nyman@linux.intel.com>
Link: https://patch.msgid.link/20260806142113.2436238-3-mathias.nyman@linux.intel.com
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
Check if the expected frame IDs for a isochronous URB submitted
mid stream is within the valid frame time window that xHC controller
is capable of queuing TDs.
The range only needs to be checked once per URB as the isoc TDs of an
URB are queued in one go with spinlock held and interrupts disabled.
Calculate the valid frame window start and end frame id in frames
instead of microframes to better match how xhci specification
section 4.11.2.5 does it.
Don't add frame id gaps or change scheduling to SIA mid stream if
the start frame is outside the valid frame winow.
Only print a debug message.
Some devices can't handle gaps in isochronous transfers.
Calculate a valid start frame for the first URB of a stream, and
align it to a full frame, or to interval start if interval is longer
than a frame
Set urb->start_frame value for every URB
cc: Dylan Robinson <dylan_robinson@motu.com>
Signed-off-by: Mathias Nyman <mathias.nyman@linux.intel.com>
Link: https://patch.msgid.link/20260806142113.2436238-2-mathias.nyman@linux.intel.com
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
ssh://gitolite.kernel.org/pub/scm/linux/kernel/git/johan/usb-serial into usb-next
Johan writes:
USB serial updates for 7.3-rc1
Here are the USB serial updates for 7.3-rc1, including:
- fix digi_acceleport port registration order
- stop digi_acceleport I/O when ports are closed
- fix digi_acceleport OOB port dev_printk()
- fix metro-usb unthrottle race
- fix option slab OOB read with malicious devices
- add support for a new class of Prolific PL256X devices
Included are also various clean ups.
All have been in linux-next with no reported issues.
* tag 'usb-serial-7.3-rc1' of ssh://gitolite.kernel.org/pub/scm/linux/kernel/git/johan/usb-serial:
USB: serial: pl2303: add support for PL256X multi-port devices
USB: serial: option: fix slab OOB read in interrupt URB callback
USB: serial: keyspan_pda: drop unused driver data usb-serial pointer
USB: serial: metro-usb: drop redundant initialisations
USB: serial: metro-usb: fix unthrottle race
USB: serial: metro-usb: replace unnecessary atomic allocation
USB: serial: digi_acceleport: fix oob port dev_printk()
USB: serial: digi_acceleport: clean up inb command submission
USB: serial: digi_acceleport: clean up write completion
USB: serial: digi_acceleport: clean up xfer buf length expression
USB: serial: digi_acceleport: drop unused in-buf define
USB: serial: digi_acceleport: stop OOB I/O when not in use
USB: serial: digi_acceleport: drop redundant driver data sanity checks
USB: serial: digi_acceleport: clean up declarations and whitespace
USB: serial: digi_acceleport: add oob port helper
USB: serial: digi_acceleport: always stop write urb on close
USB: serial: digi_acceleport: drop unused wait queue
USB: serial: digi_acceleport: fix port registration order
USB: serial: digi_acceleport: do not log stopping of urbs as errors
|
|
ssh://gitolite.kernel.org/pub/scm/linux/kernel/git/johan/usb-serial into usb-next
Johan writes:
USB serial fixes for 7.2-rc7
Here is a fix for a long-standing issue in the spcp8x5 driver which
syzbot just started hitting and a change adding lockdep annotation to
digi_acceleport to suppress a false positive deadlock warning.
Note that only the digi_acceleport commit has been in linux-next (and
with no reported issues).
* tag 'usb-serial-7.2-rc7' of ssh://gitolite.kernel.org/pub/scm/linux/kernel/git/johan/usb-serial:
USB: serial: spcp8x5: drop broken carrier detect support
USB: serial: digi_acceleport: add port lock nesting annotation
|
|
PSGMII (the Qualcomm 5-port SGMII) conveys the link negotiation result
from the PHY back to the MAC through per-channel in-band SGMII words,
exactly like SGMII and QSGMII.
However, PHY_INTERFACE_MODE_PSGMII is missing from
phylink_get_inband_type(), so phylink reports INBAND_NONE for it and
phylink_pcs_neg_mode() falls back to PHYLINK_PCS_NEG_NONE. The PCS is
then programmed in force mode and its control-register speed bits (which
default to 1000base) are used, so a slower copper link - e.g. 100base-T
- is reported as 1Gbps and cannot pass traffic.
Classify PSGMII alongside SGMII and QSGMII as INBAND_CISCO_SGMII so the
PCS negotiates in-band and the resolved link speed comes from the PHY
in-band word.
Also add PSGMII to the generic clause 22 PCS helper functions which
handle the SGMII in-band word. Without this, a PCS using these helpers
would still fall through to the default handling and force the link
state to false in phylink_mii_c22_pcs_decode_state(), fail to encode
the SGMII advertisement, and get rejected by phylink_get_link_timer_ns().
Signed-off-by: Sandeep Sondagar <sandeepsondagar@gmail.com>
Reviewed-by: Nicolai Buchwitz <nb@tipi-net.de>
Link: https://patch.msgid.link/20260809-phylink-psgmii-v3-1-908dcd3a9e3d@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
We need the USB fixes in here as well to build on top of.
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
|
|
If ->pmd_entry() sets walk->action = ACTION_AGAIN, the pmd_none() check is
retried. The PMD entry may be cleared at the point of retry.
In this case, if walk->ops->install_pte is not specified, the code
continues to the next PMD entry in the range without resetting
walk->action to ACTION_SUBTREE.
This leaves walk->action erroneously set to ACTION_AGAIN, which is
incorrect.
This was incorrect but not problematic up until commit 3b89863c3fa4
("mm/pagewalk: fix race between concurrent split and refault") which
updated walk_pud_range() to check for walk->action == ACTION_AGAIN upon
walk_pmd_range()'s return, causing the PUD walk to be retried.
In this case this results in duplicate walk callbacks being invoked,
which is erroneous and will break any caller that is not idempotent
with respect to this (and waste time for those which are). The result
is an out-of-bounds write, triggered by a local fuzzer:
[ 2.272695] ==================================================================
[ 2.273471] BUG: KASAN: slab-out-of-bounds in __mincore_unmapped_range+0x14f/0x190
[ 2.274302] Write of size 1 at addr ffff888008d9b000 by task poc/106
[ 2.274966]
[ 2.275154] CPU: 0 UID: 1000 PID: 106 Comm: poc Not tainted 7.2.0-rc6-00429-ga7c7074b58d2 #55 PREEMPT(lazy)
[ 2.275159] Hardware name: QEMU Ubuntu 24.04 PC v2 (i440FX + PIIX, arch_caps fix, 1996), BIOS 1.16.3-debian-1.16.3-2 04/01/2014
[ 2.275164] Call Trace:
[ 2.275170] <TASK>
[ 2.275172] dump_stack_lvl+0x53/0x70
[ 2.275200] print_report+0xd0/0x630
[ 2.275210] ? __pfx__raw_spin_lock_irqsave+0x10/0x10
[ 2.275219] ? irqentry_exit+0xd2/0x670
[ 2.275224] ? irqentry_exit+0xd2/0x670
[ 2.275226] ? __virt_addr_valid+0xef/0x1a0
[ 2.275239] ? __mincore_unmapped_range+0x14f/0x190
[ 2.275242] kasan_report+0xce/0x100
[ 2.275245] ? __mincore_unmapped_range+0x14f/0x190
[ 2.275248] __mincore_unmapped_range+0x14f/0x190
[ 2.275252] mincore_unmapped_range+0x45/0x70
[ 2.275254] walk_pgd_range+0xafc/0xfc0
[ 2.275261] ? __pfx_walk_pgd_range+0x10/0x10
[ 2.275264] ? __update_load_avg_se+0x3d1/0x670
[ 2.275275] __walk_page_range+0xc0/0x310
[ 2.275278] ? __pfx_find_vma+0x10/0x10
[ 2.275281] ? finish_task_switch.isra.0+0x16d/0x4f0
[ 2.275290] walk_page_range_mm_unsafe+0x26f/0x3a0
[ 2.275293] ? __pfx_mtree_load+0x10/0x10
[ 2.275298] ? __pfx_walk_page_range_mm_unsafe+0x10/0x10
[ 2.275302] ? __free_frozen_pages+0x54d/0x7e0
[ 2.275308] __do_sys_mincore+0x132/0x380
[ 2.275311] do_syscall_64+0xf9/0x540
[ 2.275316] entry_SYSCALL_64_after_hwframe+0x77/0x7f
[ 2.275322] RIP: 0033:0x422ccd
[ 2.275326] Code: b3 66 2e 0f 1f 84 00 00 00 00 00 66 90 f3 0f 1e fa 48 89 f8 48 89 f7 48 89 d6 48 89 ca 4d 89 c2 4d 89 c8 4c 8b 4c 24 08 0f 05 <48> 3d 01 f0 ff ff 73 01 c3 48 c7 c1 b8 ff ff ff f7 d8 64 89 01 48
[ 2.275329] RSP: 002b:00007fffffffec18 EFLAGS: 00000287 ORIG_RAX: 000000000000001b
[ 2.275337] RAX: ffffffffffffffda RBX: 0000000000000066 RCX: 0000000000422ccd
[ 2.275339] RDX: 00000000004d0940 RSI: 0000000001000000 RDI: 00007ffff4000000
[ 2.275340] RBP: 00000000004d0940 R08: 0000000000000100 R09: 0000000000000100
[ 2.275342] R10: 0000000000000100 R11: 0000000000000287 R12: 20c49ba5e353f7cf
[ 2.275343] R13: 00000000004990d3 R14: 0000000000000000 R15: 0000000000000001
[ 2.275346] </TASK>
[ 2.275347]
[ 2.296904] The buggy address belongs to the object at ffff888008d9b000
[ 2.296904] which belongs to the cache sigqueue of size 80
[ 2.298151] The buggy address is located 0 bytes inside of
[ 2.298151] allocated 80-byte region [ffff888008d9b000, ffff888008d9b050)
[ 2.299408]
[ 2.299601] The buggy address belongs to the physical page:
[ 2.300191] page: refcount:0 mapcount:0 mapping:0000000000000000 index:0x0 pfn:0x8d9b
[ 2.301001] flags: 0x100000000000000(node=0|zone=1)
[ 2.301535] page_type: f5(slab)
[ 2.301884] raw: 0100000000000000 ffff888107e46780 dead000000000122 0000000000000000
[ 2.302687] raw: 0000000000000000 0000000800240024 00000000f5000000 0000000000000000
[ 2.303489] page dumped because: kasan: bad access detected
[ 2.304092]
[ 2.304276] Memory state around the buggy address:
[ 2.304801] ffff888008d9af00: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
[ 2.305567] ffff888008d9af80: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
[ 2.306340] >ffff888008d9b000: fc fc fc fc fc fc fc fc fc fc fc fc fc fc fc fc
[ 2.307115] ^
[ 2.307474] ffff888008d9b080: fc fc fc fc fc fc fc fc fc fc fc fc fc fc fc fc
[ 2.308237] ffff888008d9b100: fc fc fc fc fc fc fc fc fc fc fc fc fc fc fc fc
[ 2.308997] ==================================================================
A specific example of this breaking things is mincore which walks an
internal cursor data structure a byte at a time on assumption that page
table entry callbacks are called only once for each entry.
Fix the problem by resetting walk->action to ACTION_SUBTREE prior to the
none check.
The pattern also exists in walk_pud_range() so fix it there too.
This issue was found through AI-based fuzzing.
Link: https://lore.kernel.org/20260811161949.3879321-2-imv4bel@gmail.com
Fixes: 3b89863c3fa4 ("mm/pagewalk: fix race between concurrent split and refault")
Assisted-by: Claude:claude-opus-5
Signed-off-by: Hyunwoo Kim <imv4bel@gmail.com>
Reviewed-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
Cc: Max Boone <mboone@akamai.com>
Cc: Liam R. Howlett <liam@infradead.org>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Mike Rapoport <rppt@kernel.org>
Cc: Suren Baghdasaryan <surenb@google.com>
Cc: Vlastimil Babka <vbabka@kernel.org>
Cc: <stable@vger.kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
|
|
A slot with a folio in the swap cache is freed when the folio leaves the
cache, not when its count drops. swap_put_entries_cluster() follows that
rule. swap_free_hibernation_slot() does not, it calls
__swap_cluster_free_entries() whether or not a folio sits on the slot.
Cluster readahead can put one there. It walks a raw page_cluster sized
window of offsets around the faulting entry, and a hibernation slot passes
__swap_cache_add_check() because it is not a folio and its count is not
zero. Freeing the slot then clears the entry under that folio.
The folio is now unreachable from the swap table, and the offset goes back
to the allocator. The folio is still on the LRU though, so reclaim can
pick it up later. It then takes the old offset out of folio->swap and
overwrites the table entry there, which by then may belong to someone
else.
This bug can trigger silent memory corruption, process crashes, or data
instability across completely unrelated userspace applications - typically
occurring when uswsusp is preparing the hibernation image.
I found this while working on giving hibernation slots their own marker in
the swap table, which I had discussed with Kairui.
(https://lore.kernel.org/linux-mm/abp7aDgYLrxF3Me8@KASONG-MC4/) As far as
I know there are no reports, so there is no Reported-by/Closes to add.
Check for a cached folio before freeing. The slot is then left in the
ordinary state where only the swap cache holds it, and it is freed when
the folio leaves the cache, either through the reclaim below or through
normal reclaim later.
Link: https://lore.kernel.org/20260811132209.2862708-2-youngjun.park@lge.com
Fixes: 0d6af9bcf383 ("mm, swap: use the swap table to track the swap count")
Signed-off-by: Youngjun Park <youngjun.park@lge.com>
Acked-by: Kairui Song <kasong@tencent.com>
Cc: Baoquan He <baoquan.he@linux.dev>
Cc: Barry Song <baohua@kernel.org>
Cc: Chris Li <chrisl@kernel.org>
Cc: Kemeng Shi <shikemeng@huaweicloud.com>
Cc: Nhat Pham <nphamcs@gmail.com>
Cc: <stable@vger.kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
|
|
Commit 0e2759afcaf9 ("page_counter: track failcnt only for legacy
cgroups") made failcnt accounting conditional on track_failcnt. It
enabled the flag for memcg->memory, but not for memcg->memsw or
memcg->tcpmem.
Consequently, memory.memsw.failcnt remains zero when the memory+swap limit
is hit. memory.kmem.tcp.limit_in_bytes still sets memcg->tcpmem.max, but
TCP charge failures are not reflected in memory.kmem.tcp.failcnt.
Enable failcnt accounting for both v1 counters.
To reproduce memory.memsw.failcnt:
CG=/sys/fs/cgroup/memory/memsw-test
LIMIT=33554432
mkdir "$CG"
echo "$LIMIT" > "$CG/memory.limit_in_bytes"
echo "$LIMIT" > "$CG/memory.memsw.limit_in_bytes"
Start a child process in the cgroup and make it allocate and touch 96 MiB
of memory, causing a memcg OOM.
cat "$CG/memory.memsw.failcnt"
Without the patch, memory.memsw.failcnt is 0. With the patch,
memory.memsw.failcnt is greater than 0.
To reproduce memory.kmem.tcp.failcnt:
CG=/sys/fs/cgroup/memory/tcpmem-test
LIMIT=65536
mkdir "$CG"
echo "$LIMIT" > "$CG/memory.kmem.tcp.limit_in_bytes"
Start a child process in the cgroup, create a TCP socket, and reserve
1 MiB of socket memory with SO_RESERVE_MEM. The reservation fails with
ENOMEM.
cat "$CG/memory.kmem.tcp.failcnt"
Without the patch, memory.kmem.tcp.failcnt is 0. With the patch,
memory.kmem.tcp.failcnt is greater than 0.
Link: https://lore.kernel.org/20260811030843.109104-1-guopeng.zhang@linux.dev
Closes: https://sashiko.dev/#/patchset/20260810074247.52747-1-guopeng.zhang@linux.dev?part=1
Fixes: 0e2759afcaf9 ("page_counter: track failcnt only for legacy cgroups")
Signed-off-by: Guopeng Zhang <zhangguopeng@kylinos.cn>
Acked-by: Johannes Weiner <hannes@cmpxchg.org>
Acked-by: Michal Hocko <mhocko@suse.com>
Reviewed-by: Tao Cui <cuitao@kylinos.cn>
Acked-by: Shakeel Butt <shakeel.butt@linux.dev>
Cc: Muchun Song <muchun.song@linux.dev>
Cc: Roman Gushchin <roman.gushchin@linux.dev>
Cc: <stable@vger.kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
|
|
I am seeing some rcu_tasks stalls in the Meta fleet during reclaim.
INFO: rcu_tasks detected stalls on tasks:
0000000088620d09: .. nvcsw: 6735/6735 holdout: 1 idle_cpu: -1/8
task:GlobalCPUThread state:R running task pid:2552016 tgid:2524552
Call Trace:
shrink_lruvec
mem_cgroup_iter
shrink_node
do_try_to_free_pages
try_to_free_pages
__alloc_frozen_pages_noprof
alloc_pages_noprof
pte_alloc_one
__pte_alloc
handle_mm_fault
Nothing promises direct reclaim returns in bounded time, and the scan loop
in shrink_lruvec() only calls cond_resched(), which is a no-op on
PREEMPTION kernels. Involuntary preemption is not a Tasks-RCU quiescent
state, so the reclaiming task never reports one and becomes a holdout.
Upgrade it to cond_resched_tasks_rcu_qs(), which reports a quiescent state
even when cond_resched() does nothing.
PS: This has been discussed in [1]
Link: https://lore.kernel.org/20260810-rcu_task_shrink_lruvec-v1-1-4d9f7d5251cb@debian.org
Link: https://lore.kernel.org/all/amdWVTs0WKOxguxP@gmail.com/ [1]
Signed-off-by: Breno Leitao <leitao@debian.org>
Reviewed-by: Paul E. McKenney <paulmck@kernel.org>
Acked-by: Johannes Weiner <hannes@cmpxchg.org>
Acked-by: Shakeel Butt <shakeel.butt@linux.dev>
Cc: Axel Rasmussen <axelrasmussen@google.com>
Cc: Barry Song <baohua@kernel.org>
Cc: David Hildenbrand <david@kernel.org>
Cc: Kairui Song <kasong@tencent.com>
Cc: Lorenzo Stoakes <ljs@kernel.org>
Cc: Michal Hocko <mhocko@kernel.org>
Cc: Wei Xu <weixugc@google.com>
Cc: Yuanchu Xie <yuanchu@google.com>
Cc: <stable@vger.kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
|
|
Map my old Linaro and RISCstar email addresses to my current personal
address. Neither former address receives mail anymore.
Link: https://lore.kernel.org/20260807-b4-mailmap-guodong-xu-v2-1-f7c71bc6bd9f@gmail.com
Signed-off-by: Guodong Xu <docular.xu@gmail.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
|
|
Switch to my linux.dev address and add previous one to mailmap.
Link: https://lore.kernel.org/20260807010226.8995-1-jp.kobryn@linux.dev
Signed-off-by: JP Kobryn <jp.kobryn@linux.dev>
Acked-by: Shakeel Butt <shakeel.butt@linux.dev>
Cc: Roman Gushchin <roman.gushchin@linux.dev>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
|
|
The squashfs-next.git URL hasn't been updated for many years, and it now
doesn't exist.
So remove it from the MAINTAINERS entry.
Link: https://lore.kernel.org/20260806181916.617881-1-phillip@squashfs.org.uk
Signed-off-by: Phillip Lougher <phillip@squashfs.org.uk>
Cc: Derek Barbosa <debarbos@redhat.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
|
|
memcg_reparent_objcgs() has an inherent assumption that a folio's objcg is
the objcg of the folio's node. Folio migration across nodes breaks that
assumption: the new folio simply inherits the old folio's objcg while
living on a different node.
Once the assumption is broken, the reparenting of the folio's objcg and
the reparenting of the folio's LRU list are no longer atomic.
memcg_reparent_objcgs() handles one node per iteration and drops all the
locks in between, so the objcg gets reparented in the iteration for the
objcg's node while the LRU list gets spliced in the iteration for the
folio's node. Any LRU operation on that folio in between resolves its
lruvec through the objcg, and thus takes the lru_lock of the wrong memcg,
not the lru_lock of the list the folio is actually on.
Fix this by selecting the objcg by folio_nid() at charge time, and by
re-deriving it for the destination node in mem_cgroup_migrate() and
mem_cgroup_replace_folio().
Link: https://lore.kernel.org/20260807142406.443516-1-shakeel.butt@linux.dev
Fixes: f1cf8d2f36dc ("mm: memcontrol: eliminate the problem of dying memory cgroup for LRU folios")
Signed-off-by: Johannes Weiner <hannes@cmpxchg.org>
Signed-off-by: Shakeel Butt <shakeel.butt@linux.dev>
Reported-by: Karl Erik Hofseth <karl.e.hofseth@opoint.com>
Closes: https://lore.kernel.org/all/anMmd1ADrDVwMO6v@work/
Co-developed-by: Johannes Weiner <hannes@cmpxchg.org>
Acked-by: Muchun Song <muchun.song@linux.dev>
Acked-by: Qi Zheng <qi.zheng@linux.dev>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Roman Gushchin <roman.gushchin@linux.dev>
Cc: <stable@vger.kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
|
|
The ffs()/fls() guard in ethnl_set_tsconfig() was meant to enforce
that the user selects exactly one tx_type (and one rx_filter)
at a time (off / none are explicit types with non-zero values).
However, both ffs(0) and fls(0) return 0, so the guard passes
a zero-valued bitset through.
The subsequent ffs(req_tx_type) - 1 would produce -1, if user selected
no bit. net_hwtstamp_validate() catches the invalid -1 downstream,
but returns a generic error (-ERANGE) without telling the user
what went wrong. Return -EINVAL + extack instead.
Replace the ffs()/fls() comparison with a hweight32() == 1 check.
Reviewed-by: Andrew Lunn <andrew@lunn.ch>
Reviewed-by: Joe Damato <joe@dama.to>
Reviewed-by: Vadim Fedorenko <vadim.fedorenko@linux.dev>
Link: https://patch.msgid.link/20260812162230.1837788-1-kuba@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
The IRQ vectors allocated in stmmac_config_multi_msi() or
stmmac_config_single_msi() where never explicitly cleaned up. As
pcim_enable_device() is used, all sorts of other functions are switched
to managed mode. The missing cleanup here isn't actually missing, it's
buried in the depths of PCI code.
But: There are some ongoing activities to remove that cleanup magic.
See the linked discussions below.
This patch prepares the dwmac-intel code for the removal.
Link: https://lore.kernel.org/netdev/27fec7d0ed633218a7787be3edce63c3038c63e2.camel@mailbox.org/
Link: https://lore.kernel.org/netdev/7e024db2557a4d5822a0dd409ae678d10d815d9c.camel@mailbox.org/
Signed-off-by: Florian Bezdeka <florian.bezdeka@siemens.com>
Link: https://patch.msgid.link/20260810-flo-net-stmmac-default-affinity-core-v2-1-d2105780b8ca@siemens.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
This debugfs file isn't used by kernel's selftests, so drop it.
Reported-by: syzbot+3147c5de186107ffc7a1@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=3147c5de186107ffc7a1
Suggested-by: Jakub Kicinski <kuba@kernel.org>
Signed-off-by: Slawomir Stepien <sst@poczta.fm>
Link: https://patch.msgid.link/20260810085717.570382-1-sst@poczta.fm
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
tun_get_user() uses tun->align both as skb headroom and when choosing how
much packet data to keep linear. OVS can propagate an oversized headroom
request from another port to TUN or TAP.
When align is larger than the usable space in a one-page skb head,
SKB_MAX_HEAD(align) underflows and the result becomes negative when stored
in good_linear. That value later wraps when assigned to the size_t linear
variable, and tun_alloc_skb() can place skb->data outside the allocated
head.
Bound the headroom stored by TUN to the one-page skb-head budget and the
largest non-sentinel 16-bit skb header offset. Leave one linear byte for
raw TUN and a complete Ethernet header for TAP, including NET_IP_ALIGN.
Also pull the raw-TUN protocol byte and the TAP Ethernet header before
accessing them, so these checks remain safe for nonlinear skbs supplied by
other allocation paths.
Fixes: eaea34b23c46 ("net/tun: implement ndo_set_rx_headroom")
Cc: stable@vger.kernel.org
Signed-off-by: Asim Viladi Oglu Manizada <manizada@pm.me>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260812012139.2134643-1-manizada@pm.me
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
l2tp_tunnel_notify() and l2tp_session_notify() use
genlmsg_multicast_allns(), which delivers to listeners in every network
namespace. l2tp is per-namespace, and a tunnel records the namespace it
belongs to in tunnel->l2tp_net. Each event concerns one namespace, yet
every namespace is told about it. A tunnel event carries the tunnel and
peer tunnel ids, plus the socket's addresses with both ports for a UDP
tunnel. A session event carries the session and peer session ids, the
interface name, plus the L2TP cookies where those are set. A listener
needs no privilege for any of this, because l2tp_multicast_group[]
carries no flags and genl_bind() asks for no capability.
The fix is to send to the tunnel's namespace with
genlmsg_multicast_netns(). Commit 134e63756d5f ("genetlink: make netns
aware") added both helpers and drew the line between them. The netns
variant is for an object that lives in a namespace.
I found this by auditing the tree's six genlmsg_multicast_allns() call
sites for objects that live in a network namespace. Only the two l2tp
ones do.
I reproduced it on net at dd057113ac7b, in a virtual machine, with no
real hardware involved. A process in the initial namespace, running as
an ordinary user with an empty capability set, receives the create and
delete events of a tunnel. The tunnel was set up inside an unprivileged
user and network namespace. tools/testing/selftests/net/l2tp.sh passes
before and after.
On a container host, any local user and every other tenant can read a
tenant's tunnel parameters.
Cc: stable+noautosel@kernel.org # high regression risk
Signed-off-by: Maoyi Xie <maoyixie.tju@gmail.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260809094252.2107242-1-maoyixie.tju@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
'net-smc-close-the-smc-d-teardown-window-around-the-ghost-send-buffer'
Bryam Vargas says:
====================
net/smc: close the SMC-D teardown window around the ghost send buffer
Both patches only matter on the SMC-D DMB-nocopy path, where the ghost send
buffer exists, and the only in-tree provider of support_mmapped_rdmb is
dibs_loopback. CONFIG_DIBS_LO is default n and its help calls it a testing aid,
so on a stock config neither bug is reachable.
v1 moved smcd_buf_detach() after the drain. Dust Li replied that it does not
fully eliminate the race and asked whether RCU is the better shape. He is right
about the first part; I built both and measured them.
An SMC-D loopback KASAN rig, one module binary, teardown form selected at
runtime. "path" counts connections reaching either teardown site with the link
group already unlinked, "armable" how many of those still had both gates in
smcd_handle_irq() open when the drain returned, "re-armed" the device arming
the tasklet again afterwards:
form path armable re-armed
upstream 169 73 29
v1 (drain, then detach) 172 78 33
unregister first, then drain 31 0 0
v1 + RCU 24 9 3
Two caveats on that table. The last two arms ran far shorter than the first two,
so compare the armable/re-armed ratios rather than the absolute path counts. And
the third row also forced tasklet_kill() in the !soft path, which 1/2 does not;
that was inert here because smc_lgr_terminate_work() passes soft=true, so the
same call ran either way.
The reorder alone leaves the window open, which is what Dust saw. RCU doesn't
close it either: smcd_buf_detach() both frees the descriptor and clears the
field, and RCU defers only the free, so a re-armed tasklet still runs and still
finds conn->sndbuf_desc NULL. The gate has to be shut before the drain, and
that is 1/2. RCU on the descriptor would still be a reasonable thing to want
for the free itself; it just doesn't substitute for 1/2, so I didn't fold it
in. Your call if you want it anyway.
Caveat on 1/2: the two changes the table covers -- unconditional
smc_ism_unset_conn(), and drain before detach -- were measured together, not
separately. It also clears conn->sndbuf_desc before freeing it, so a reader
that samples the pointer cannot get one that is already freed; that part is
by inspection.
2/2 is a second dereference the same teardown reaches, found while running the
above. smc_close_stream_wait() calls smc_tx_prepared_sends() from inside
sk_wait_event(), which evaluates its condition once with the socket lock
released, and a terminating link group clears conn->sndbuf_desc right there.
SIOCOUTQ reads the same field by hand, and smc_close_cancel_work() drops the
socket lock across two cancel_*_sync() calls, so 2/2 bounds that too. Eight
faults across three boots, the earliest 89 seconds in:
RIP: smc_close_stream_wait+0x66d [smc]
smc_close_active -> __smc_release -> smc_release -> __x64_sys_close
The faulting address is NULL plus offsetof(struct smc_buf_desc, len), nothing
there is attacker-chosen, and the value read never reaches userspace, so there
is no memory-safety primitive and no leak oracle -- it is an oops. The task dies
inside close() holding the socket lock, so I would expect the socket to leak with
it, but I didn't isolate that from the rig's own effects and I'm not claiming it.
Reaching either bug needs a link-group teardown while a socket is parked in that
wait. smc_lgr_cleanup_early() off a failed first-contact handshake gets there, as
does smc_clc_wait_msg() on a peer DECLINE with FIRST_CONTACT -- both by
inspection. The rig instead drove smc_lgr_terminate_sched() from a debug module
parameter, so only the initiation is synthetic; the unlink, the deferred worker,
smc_conn_free() and smcd_handle_irq() are the unmodified path. Logs and the rig
on request.
I haven't touched tasklet_unlock_wait() in the !soft path of smc_conn_kill().
It waits out TASKLET_STATE_RUN without clearing TASKLET_STATE_SCHED, but I have
no measurement showing that reachable here, so it stays as it is.
====================
Link: https://patch.msgid.link/20260808-b4-disp-22f119e6-v2-0-61647601a6f3@proton.me
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
smc_close_stream_wait() calls smc_tx_prepared_sends() from inside its
sk_wait_event() condition, and sk_wait_event() evaluates that condition
once with the socket lock released. smcd_buf_detach() clears
conn->sndbuf_desc from smc_conn_kill() under lock_sock(), so a link group
terminating while a socket waits there leaves the helper dereferencing
NULL, faulting out of close(). SIOCOUTQ reads the field by hand, and
smc_close_cancel_work() drops the lock across two cancel_*_sync() calls.
Sample the pointer once in the helper, report nothing prepared while it is
unset, and bound the ioctl the same way. The receive tasklet dereferences
the field directly in smc_cdc_msg_recv_action(), not through this helper;
1/2 is what keeps it from running that late.
Fixes: ae2be35cbed2 ("net/smc: {at|de}tach sndbuf to peer DMB if supported")
Cc: stable@vger.kernel.org
Signed-off-by: Bryam Vargas <hexlabsecurity@proton.me>
Reviewed-by: Sidraya Jayagond <sidraya@linux.ibm.com>
Reviewed-by: Tony Lu <tonylu@linux.alibaba.com>
Link: https://patch.msgid.link/20260808-b4-disp-22f119e6-v2-2-61647601a6f3@proton.me
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
smc_conn_free() calls smc_ism_unset_conn() only while the link group is
still on its device list, and never sets conn->killed.
smc_lgr_terminate_sched() unlinks the group immediately and defers killing
its connections to a work item, so a connection freed in that window keeps
its smcd->conn[] slot with both gates in smcd_handle_irq() open, and the
device can re-arm the receive tasklet after tasklet_kill() has returned. On
the DMB-nocopy path the ghost send buffer is freed right after that drain,
so the re-armed tasklet dereferences it.
Unregister unconditionally and drain before the detach at both teardown
sites, mirroring rmb_desc, which smc_buf_unuse() releases after the drain.
Clear conn->sndbuf_desc before freeing it as well, so a reader that samples
the pointer cannot get one that is already freed.
Fixes: ae2be35cbed2 ("net/smc: {at|de}tach sndbuf to peer DMB if supported")
Cc: stable@vger.kernel.org
Signed-off-by: Bryam Vargas <hexlabsecurity@proton.me>
Reviewed-by: Sidraya Jayagond <sidraya@linux.ibm.com>
Reviewed-by: Tony Lu <tonylu@linux.alibaba.com>
Link: https://patch.msgid.link/20260808-b4-disp-22f119e6-v2-1-61647601a6f3@proton.me
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Xuanqiang Luo says:
====================
net: phy: dp83640: fix shared clock lifetime and probe error cleanup
The DP83640 driver shares one PTP clock between all PHYs on the same MII
bus.
Its driver-local clock lookup and removal scheme can leak the shared clock
on probe failure or free it while another probe is acquiring it.
This series moves the shared clock to the PHY package infrastructure.
Patch 1 adds PHY package locking helpers.
Patch 2 embeds the pin configuration in the shared clock.
Patch 3 clears per-PHY state when PTP clock registration fails.
Patch 4 fixes the shared clock lifetime using the PHY package
infrastructure.
====================
Link: https://patch.msgid.link/20260811151345.73582-1-xuanqiang.luo@linux.dev
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
Commit 42e2a9e11a1d ("net: phy: dp83640: improve phydev and driver
removal handling") moved per-bus clock cleanup from module exit to the
remove path. This leaves two lifetime problems.
dp83640_clock_get_bus() publishes a newly allocated clock before the
driver allocates its per-PHY data and registers the PTP clock. If either
operation fails, no PHY is bound and the remove callback cannot release
the clock, leaking the clock and the MII bus device reference.
The remove path can also free a clock after dropping clock_lock. A
concurrent probe may already have found the clock under
phyter_clocks_lock and be waiting for clock_lock, allowing it to acquire
a freed mutex and access the freed clock.
Use the PHY package infrastructure for the per-bus clock. PHY packages
are tracked per MII bus, and the driver uses BROADCAST_ADDR as the
package key so the DP83640 PHYs on the same bus share the same clock
storage. Call phy_package_join() during probe and phy_package_leave() on
probe errors and in remove.
Serialize the one-time clock initialization with the package lock because
phy_package_probe_once() elects an initializer but does not wait for
initialization to finish.
Cc: stable+noautosel@kernel.org # untested fix to a driver init path
Reviewed-by: Andrew Lunn <andrew@lunn.ch>
Signed-off-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
Link: https://patch.msgid.link/20260811151345.73582-5-xuanqiang.luo@linux.dev
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
dp83640_probe() publishes its per-PHY state through phydev before
registering the PTP clock. If registration fails, the private data is
freed while phydev->mii_ts and phydev->priv still point to it, and
default_timestamp remains set.
Clear the published PHY state and reset the PTP clock pointer before
freeing the private data.
Cc: stable+noautosel@kernel.org # untested fix to a driver init path
Reviewed-by: Andrew Lunn <andrew@lunn.ch>
Signed-off-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
Link: https://patch.msgid.link/20260811151345.73582-4-xuanqiang.luo@linux.dev
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
The DP83640 has a fixed number of PTP pins, and its pin configuration
has the same lifetime as the per-bus clock. Allocating the configuration
separately adds an allocation failure path and requires a separate free.
Embed the pin configuration in struct dp83640_clock and point the PTP
clock information at the embedded array. This changes only the storage;
the pin functions remain configurable at runtime. It also allows all
per-bus clock storage to be managed as one allocation.
Reviewed-by: Andrew Lunn <andrew@lunn.ch>
Signed-off-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
Link: https://patch.msgid.link/20260811151345.73582-3-xuanqiang.luo@linux.dev
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
The PHY package API provides private data shared by all PHYs in a
package. Drivers are responsible for synchronizing access to this data,
but the API does not provide a lock for that purpose.
Add phy_package_lock() and phy_package_unlock() for drivers to serialize
access to package-private data, including its initialization.
Reviewed-by: Andrew Lunn <andrew@lunn.ch>
Signed-off-by: Xuanqiang Luo <luoxuanqiang@kylinos.cn>
Link: https://patch.msgid.link/20260811151345.73582-2-xuanqiang.luo@linux.dev
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
|
|
git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace
Pull tracing fixes from Steven Rostedt:
- Fix NULL pointer dereference when matching unloaded module wildcard
event
The set_event can take events for modules that have not been loaded
yet. This is done by writing '<event>:mod:<module>'.
If '<event>' is not added, then it means to add all events in
<module>. This wildcard is represented by a NULL pointer. If one were
to try to remove the same module item with a named event it would
cause a NULL pointer dereference when comparing the NULL with the
name in strcmp().
echo ':mod:kvm' > /sys/kernel/tracing/set_event
echo '!kvm_ack_irq:mod:kvm' >> /sys/kernel/tracing/set_event
The above will do a strcmp("kvm_ack_irq", NULL) and crash the kernel.
Test for NULL (wildcard) before doing the strcmp().
- Fix event data field race in loading two modules at the same time
When a module loads, its trace events get registered. The fields of
the events are also dynamically created and added to the events
fields list. It also will call a function that will look at all the
events for updates that need to be done. If two modules load at the
same time, the one that scans all events and their fields may read
the one being added as the scan doesn't take the event_mutex. This
may cause a data race.
Have the scan take the event_mutex to prevent the race.
* tag 'trace-v7.2-rc7' of git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace:
tracing: Fix race between update_event_fields and, event_define_fields
tracing: Fix NULL pointer dereference in module event cache removal
|
|
The THIS_MODULE series [1] applied to rust-next moved ThisModule from
lib.rs into a module.rs submodule, making the tuple struct field private
outside the module. This breaks the module.0 field access in serdev in
driver-core-next.
Update the call to __serdev_device_driver_register() to use the public
module.as_ptr() accessor to fix the build.
Link: https://lore.kernel.org/all/20260811-fix-fops-owner-v10-0-7e71776f9dbe@linux.dev/ [1]
Closes: https://lore.kernel.org/all/DKNAS52KYWLD.M15VEC6U0F6R@kernel.org/
Reviewed-by: Gary Guo <gary@garyguo.net>
Reviewed-by: Markus Probst <markus.probst@posteo.de>
Link: https://patch.msgid.link/20260813152444.514580-1-dakr@kernel.org
Signed-off-by: Danilo Krummrich <dakr@kernel.org>
|
|
Eduard Zingerman says:
====================
selftests/bpf: fix for veristat file/prog filters processing
At the moment veristat filtering behaves unexpectedly for the
following filter expression:
-f !file/prog
The expression rejects all programs with name 'prog', and all programs
in a file with name 'file'. Fix the expression to exclude only a
program 'prog' from a file 'file', also add a set of tests to exercise
filtering logic.
Changelog:
v1 -> v2:
- added fixes tag for patch #1 (bot+bpf-ci);
- extended test cases for '!*foo*' and '*foo*' filters in patch #2
(bot+bpf-ci);
- added patch #3, replacing direct read() calls with calls to
read_output(), guaranteeing input buffer null termination
(bot+bpf-ci).
v1: https://lore.kernel.org/bpf/20260811-veristat-filter-fix-v1-0-b5b43c431550@gmail.com/
---
====================
Link: https://patch.msgid.link/20260811-veristat-filter-fix-v2-0-6c234c4cd6ef@gmail.com
Signed-off-by: Andrii Nakryiko <andrii@kernel.org>
|
|
In veristat tests replace direct read() calls with calls to
read_output() utility function, which:
- guarantees that the input buffer is zero terminated;
- asserts that read operation succeeded.
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
Signed-off-by: Andrii Nakryiko <andrii@kernel.org>
Link: https://lore.kernel.org/bpf/20260811-veristat-filter-fix-v2-3-6c234c4cd6ef@gmail.com
|
|
Test cases for veristat file/prog name filtering logic.
Check various formulations for any (*foo*), file (*foo*/),
prog (/bar) and file/prog (*foo*/bar) filters, alongside
erroneous filters and mixed allow/deny filter expressions.
Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>
Signed-off-by: Andrii Nakryiko <andrii@kernel.org>
Link: https://lore.kernel.org/bpf/20260811-veristat-filter-fix-v2-2-6c234c4cd6ef@gmail.com
|