summaryrefslogtreecommitdiff
AgeCommit message (Collapse)Author
2026-06-30sched/debug: Remove unused schedstatsShrikanth Hegde
nr_migrations_cold, nr_wakeups_passive and nr_wakeups_idle are not being updated anywhere. So remove them. These are per process stats. So updating sched stats version isn't necessary. Signed-off-by: Shrikanth Hegde <sshegde@linux.ibm.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Reviewed-by: K Prateek Nayak <kprateek.nayak@amd.com> Tested-by: K Prateek Nayak <kprateek.nayak@amd.com> Link: https://patch.msgid.link/20260625124648.802832-2-sshegde@linux.ibm.com
2026-06-30sched/psi: skip irqtime accounting when no new irq time has elapsedUsama Arif
psi_account_irqtime() reads irq_time_read() into a per-rq cumulative counter and only bails out when the delta vs. the previously accounted amount is negative. A delta of exactly zero is treated as "do the work": psi_write_begin() is taken, cpu_clock(cpu) is read (which on x86 ends up in native_sched_clock() / rdtsc) and the cgroup ancestor chain is walked to add zero to every group's PSI_IRQ_FULL bucket. The zero-delta case is common in practice -- it fires every time a context switch crosses a PSI group boundary on a CPU that hasn't serviced an interrupt between the two switches. Measured on a 176-thread AMD EPYC 9D64 server running a compute intensive production workload, instrumented with bpftrace over a 30s window (irq_time_read() read directly from the per-CPU cpu_irqtime so that delta == 0 and delta < 0 could be separated): @total 17,229,311 (100.0%) @ret_curr_swapper 7,864,195 ( 45.6%) curr->pid == 0 @ret_samegrp 323,299 ( 1.9%) same cgroup as prev @reached_delta 9,041,817 ( 52.5%) @delta_positive 6,358,192 ( 36.9%) real work @delta_zero 2,683,625 ( 15.6%) work wasted (this patch) @delta_negative (0) ( 0.0%) monotonic clock So 15.6 % of all psi_account_irqtime() calls - and 29.7 % of the calls that get past the early returns - hit the delta == 0 case; delta < 0 did not occur once in the 30 s window. Under the current code each of those ~89 k calls per second performs the full seqcount write + cpu_clock() read + cgroup-chain walk just to add 0 to every group's PSI_IRQ_FULL counter. Extend the early-return to also cover delta == 0. rq->psi_irq_time does not need updating in that case (it would store the same value back) and no PSI bucket would change. The existing behaviour for delta > 0 is untouched. Signed-off-by: Usama Arif <usama.arif@linux.dev> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Reviewed-by: Shakeel Butt <shakeel.butt@linux.dev> Acked-by: Johannes Weiner <hannes@cmpxchg.org> Link: https://patch.msgid.link/20260617175219.2494857-2-usama.arif@linux.dev
2026-06-30sched/fair: Reflow sched_balance_rq()Peter Zijlstra
Reflow to reduce indenting. Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Link: https://patch.msgid.link/20260618105627.GP49951@noisy.programming.kicks-ass.net
2026-06-30sched/fair: Simplify balance_interval reset logic in sched_balance_rq()Xin Zhao
Because active_balance is initialized to 0, and need_active_balance() is a pre-condition for setting it to 1, the condition '!active_balance || need_active_balance()' is a truism and can be removed. Signed-off-by: Xin Zhao <jackzxcui1989@163.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Link: https://patch.msgid.link/20260617072151.1173416-3-jackzxcui1989@163.com
2026-06-30sched/fair: Don't trigger active lb if src_rq->curr is not on_rqXin Zhao
Active load balancing relies on migration threads, which temporarily preempt tasks on the source runqueue (src_rq). This preemption can negatively impact overall system performance. The active balancing logic includes a check to verify whether the current task (curr) on src_rq can actually run on the destination runqueue (dst_rq). We have observed that when curr is a CFS task and its on_rq flag is 0, the active balancing failure rate is exceptionally high. The following table summarizes test data collected over 300 seconds on an 18-CPU platform under a specific fillback task scenario: fair: busiest->curr->sched_class == &fair_sched_class on_rq: busiest->curr->on_rq total: active balance count triggered of correspondent type fail: fail to migrate one task in active_load_balance_cpu_stop() fair && !on_rq !fair && !on_rq domain total fail total fail cpu0 0x00003 0 0 0 0 cpu0 0x3ffff 33 33 1 1 cpu1 0x00003 0 0 0 0 cpu1 0x3ffff 42 42 0 0 cpu2 0x0003c 4 4 0 0 cpu2 0x3ffff 12 12 0 0 cpu3 0x0003c 3 3 0 0 cpu3 0x3ffff 8 7 0 0 cpu4 0x0003c 2 2 0 0 cpu4 0x3ffff 5 4 0 0 cpu5 0x0003c 4 4 0 0 cpu5 0x3ffff 8 8 0 0 cpu6 0x003c0 60 60 0 0 cpu6 0x3ffff 28 27 0 0 cpu7 0x003c0 194 184 0 0 cpu7 0x3ffff 35 35 1 1 cpu8 0x003c0 240 228 0 0 cpu8 0x3ffff 28 28 0 0 cpu9 0x003c0 0 0 0 0 cpu9 0x3ffff 10 10 0 0 cpu10 0x03c00 52 50 0 0 cpu10 0x3ffff 0 0 0 0 cpu11 0x03c00 70 68 0 0 cpu11 0x3ffff 1 1 0 0 cpu12 0x03c00 73 72 0 0 cpu12 0x3ffff 0 0 0 0 cpu13 0x03c00 79 76 0 0 cpu13 0x3ffff 0 0 0 0 cpu14 0x3c000 0 0 0 0 cpu14 0x3ffff 57 55 1 0 cpu15 0x3c000 53 52 1 0 cpu15 0x3ffff 30 29 0 0 cpu16 0x3c000 344 341 10 6 cpu16 0x3ffff 103 100 2 1 cpu17 0x3c000 183 179 2 2 cpu17 0x3ffff 78 77 0 0 sum 1839 1791 18 11 In __schedule(), before curr is updated to next, pick_next_task() invokes sched_balance_rq(). This function temporarily unlocks and relocks the runqueue, creating a window where other CPUs may observe rq->curr->on_rq as 0. We can safely skip active balancing when src_rq->curr->on_rq == 0, as other eligible tasks have likely already been evaluated. We retain the affinity check on dst_rq to trigger active balancing, since such tasks are often woken by (or wake up) tasks on src_rq that share similar affinity constraints. Furthermore, detach_tasks() releases the runqueue lock; any tasks awakened during this window may preempt the previous CFS task. My testing (data not shown) indicates that active balancing succeeds in 98.4% of cases where !fair && on_rq. This scenario does not require a stop-work callback, but would necessitate an additional detach/attach path. As Valentin and Vincent have already discussed, this addition does not appear justified at this time (see [1]). Since can_migrate_task() already checks on_cpu during the cfs_tasks traversal, adding an on_rq check will have negligible performance overhead due to cache locality. There are two reasons for not combining the on_rq check with the cpumask_test_cpu() check: - Avoiding new scenarios that would skip the logic for resetting balance_interval to min_interval. - The existing check for whether the busiest CPU recently triggered active load balancing already filters more cases than the on_rq check. [1]: https://lore.kernel.org/lkml/20190815145107.5318-5-valentin.schneider@arm.com/ Signed-off-by: Xin Zhao <jackzxcui1989@163.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Reviewed-by: Valentin Schneider <vschneid@redhat.com> Link: https://patch.msgid.link/20260617072151.1173416-2-jackzxcui1989@163.com
2026-06-30sched/eevdf: Speedup short slice task schedulingVincent Guittot
When a task with a shorter slice is enqueued, we protect the running task which has a longer slice until it becomes ineligible instead of a full slice in order to speedup the switch to other tasks until the task with the shortest slice is scheduled. This helps to the task to not wait too many full slices before running. Signed-off-by: Vincent Guittot <vincent.guittot@linaro.org> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Tested-by: K Prateek Nayak <kprateek.nayak@amd.com> Link: https://patch.msgid.link/20260624151229.1710703-7-vincent.guittot@linaro.org
2026-06-30sched/eevdf: Always update slice protectionVincent Guittot
Even if p will not preempt current, it modifies the avg_vruntime and possibly the min slice. Make sure to update the slice protection with the updated figures. As an example, Batch and Sched Idle tasks can otherwise get a larger lag than their slice and finaly delay the scheduling of a normal task, which deadline will be a later. Signed-off-by: Vincent Guittot <vincent.guittot@linaro.org> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Tested-by: K Prateek Nayak <kprateek.nayak@amd.com> Link: https://patch.msgid.link/20260624151229.1710703-6-vincent.guittot@linaro.org
2026-06-30sched/eevdf: Cancel slice protection if short slice task is eligibleVincent Guittot
If a short slice task will not be the next to be picked but is eligible, we cancel the slice protection to speedup the time when the short slice task will be the next to run. Signed-off-by: Vincent Guittot <vincent.guittot@linaro.org> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Tested-by: K Prateek Nayak <kprateek.nayak@amd.com> Link: https://patch.msgid.link/20260624151229.1710703-5-vincent.guittot@linaro.org
2026-06-30sched/eevdf: Update slice protection even when resched is already setVincent Guittot
Even if resched is already set, we might want to update or even cancel the slice protection and ensure that the newly waking task will be the next one to run. Signed-off-by: Vincent Guittot <vincent.guittot@linaro.org> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Tested-by: K Prateek Nayak <kprateek.nayak@amd.com> Link: https://patch.msgid.link/20260624151229.1710703-4-vincent.guittot@linaro.org
2026-06-30sched/eevdf: Take into account current's lag when updating slice protectionVincent Guittot
Take into account the lag of current task when updating the slice protection in order to ensure that the absolute value of lags will remain in the range [0 : slice+tick] A task that already has a negative lag will see its protection reduced whereas a task with positive lag will keep a full slice protection. Signed-off-by: Vincent Guittot <vincent.guittot@linaro.org> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Tested-by: K Prateek Nayak <kprateek.nayak@amd.com> Link: https://patch.msgid.link/20260624151229.1710703-3-vincent.guittot@linaro.org
2026-06-30sched/fair: Set next buddy for preempt shortVincent Guittot
If a shorter slice task can preempt current at wakeup, we make sure that the decision will not be overwritten in between by setting the task as the next buddy. This still implies that the waking task remains eligible when the scheduler will actually pick the next task to run. Signed-off-by: Vincent Guittot <vincent.guittot@linaro.org> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Tested-by: K Prateek Nayak <kprateek.nayak@amd.com> Link: https://patch.msgid.link/20260624151229.1710703-2-vincent.guittot@linaro.org
2026-06-30sched/fair: Remove unused arguments from set_preempt_buddy()K Prateek Nayak
On a tangential note, I just noticed set_preempt_buddy() has two unused parameters. Seems to have been like that since it was introduced in commit e837456fdca8 ("sched/fair: Reimplement NEXT_BUDDY to align with EEVDF goals"). Signed-off-by: K Prateek Nayak <kprateek.nayak@amd.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
2026-06-30sched/eevdf: Move to a single runqueuePeter Zijlstra (Intel)
Change fair/cgroup to a single runqueue. Infamously fair/cgroup isn't working for a number of people; typically the complaint is latencies and/or overhead. The latency issue is due to the intermediate entries that represent a combination of tasks and thereby obfuscate the runnability of tasks. The approach here is to leave the cgroup hierarchy as is; including the intermediate enqueue/dequeue but move the actual EEVDF runqueue outside. This means things like the shares_weight approximation are fully preserved. That is, given a hierarchy like: R | se--G1 / \ G2--se se--G3 / \ | T1--se se--T2 se--T3 This is fully maintained for load tracking, however the EEVDF parts of cfs_rq/se go unused for the intermediates and are instead connected like: _R_ / | \ T1 T2 T3 Since the effective weight of the entities is determined by the hierarchy, this gets recomputed on enqueue,set_next_task and tick. Notably, the effective weight (se->h_load) is computed from the hierarchical fraction: se->load / cfs_rq->load. Since EEVDF is now exclusively operating on rq->cfs, it needs to consider cfs_rq->h_nr_queued rather than cfs_rq->nr_queued. Similarly, only tasks can get delayed, simplifying some of the cgroup cleanup. One place where additional information was required was set_next_task() / put_prev_task(), where we need to track 'current' both in the hierarchical sense (cfs_rq->h_curr) and in the flat sense (cfs_rq->curr). As a result of only having a single level to pick from, much of the complications in pick_next_task() and preemption go away. Since many of the hierarchical operations are still there, this won't immediately fix the performance issues, but hopefully it will fix some of the latency issues. TODO: split struct cfs_rq / struct sched_entity TODO: try and get rid of h_curr Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Link: https://patch.msgid.link/20260605124052.227463677%40infradead.org
2026-06-30sched/fair: Change the default cgroup_mode to concurPeter Zijlstra
For all the reasons described in the preceding patches, the way cgroup weight is computed is problematic. However, changing it is bound to also lead to trouble. Esp. since people might have taken to inflating the weight value where they can. Since things are configurable, change the default and hope this serves more people than it hurts, esp. in the longer run. Specifically, this prepares for a flattened runqueue, where the hierarchical weight becomes far more important (F_g^d terms), so getting rid of small F_g is imperative. Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Link: https://patch.msgid.link/20260605124052.080482755%40infradead.org
2026-06-30sched/fair: Add cgroup_mode: tasksPeter Zijlstra
Since we are exploring this space; include a scheme that scales by total number of runnable tasks. This results in: F_g_n' = M * F_g_n This will obviously have: avg(F_g_n') > 1, (it will be ~M/N in fact). And while that sounds odd, it actually has a fairly straight foward meaning for "cpu.weight": average weight per member task. This is an entirely valid and workable option, it is however wildly different from the traditional meaning. Included for completeness (and curiosity). Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Link: https://patch.msgid.link/20260605124051.921991975%40infradead.org
2026-06-30sched/fair: Add cgroup_mode: concurPeter Zijlstra
Improve upon the previous scheme ("max") by no longer assuming maximal concurrency. Instead scale by: 'min(nr_tasks, nr_cpus)'. This handles the low concurrency cases more gracefully: F_g_n' = min(M, N) * F_g_n Notably this is the first mode where: avg(F_g_n) = 1 In the single task case it reduces to ("smp") and then it nicely scales up until it hits N, where it behaves like ("max"). This is no longer clipped at nice -20. Strictly speaking it isn't different from the normal SMP scenario where all tasks are extremely unbalanced. There are no unnatural inflations in this scheme. The meaning of "cpu.weight" would be: weight per active CPU. NOTE: Compute the group wide number of tasks by extending the tg->load_avg computation with tg->runnable_avg, since cfs_rq->runnable_avg is based on cfs_rq->h_nr_running. Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Link: https://patch.msgid.link/20260605124051.740585993%40infradead.org
2026-06-30sched/fair: Add cgroup_mode: maxPeter Zijlstra
In order to avoid the average CPU fraction avg(F_g_n) becoming tiny '1/N', assume each cgroup is maximally concurrent and distrubute 'N*weight', such that: F_g_n' = N * F_g_n Giving: avg(F_g_n') = N*avg(F_g_n) ~ N * 1/N = 1 And while this sounds like it solves things, remember what that ~ meant. There is the corner case when a cgroup is minimally loaded, eg a single runnable task, therefore limit the CPU fraction to that of a nice -20 task to avoid getting too much load. This last bit is what makes it different from a previous proposal to allow raising cpu.weight to '100 * N', that would not limit the mininal concurrency case and results in a very large F_g_n. And just like F_g_n << 1 is problematic, so is F_g_n >> 1 for the exact same reasons (it would drown the kthreads, but it also risks overflowing the load values). So while this might appear to be a better scheme than the current default scheme, it doesn't really handle less than maximal concurrency nicely -- it clips and introduces artificially large weights. So where the traditional SMP mode works well when nr_tasks << nr_cpus, MAX doesn't work well in that regime and vice-versa. The meaning of "cpu.weight" would be: weight per allowed CPU. Included for completeness (and infrastructure). Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Link: https://patch.msgid.link/20260605124051.589618504%40infradead.org
2026-06-30sched/fair: Add cgroup_mode: upPeter Zijlstra
Instead of calculating the proportional fraction of the group weight for each CPU, just give each CPU the full measure, ignoring these pesky SMP problems. This makes the SMP cgroup fraction (F_g_n) equal to 1, and ensures a single task in a cgroup competes on equal footing to a task in a level above. However, as already explored, this is not a very good policy because it gets the SMP weight distribution wrong. Included for completeness. Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Link: https://patch.msgid.link/20260605124051.450303977%40infradead.org
2026-06-30sched/fair: Add cgroup_mode switchPeter Zijlstra
The effective task weight (W_t') for a task in cgroup g on CPU n is given by: W_t W_t' = W_g * F_g_n * ---------- \Sum W_t_n Where W_g is the group's weight (cpu.weight), F_g_n is the fraction of the group weight for CPU n and W_t/W is the relative weight of this task against all other tasks in the same group on the same CPU. Furthermore, this makes: \Sum W_t_n F_g_n = ---------- \Sum W_t The fraction of weight inside the group of CPU n against the whole group. The problem is with F_g_n, the primary goal of this fraction is to make sure that the relative weight of tasks, when distributed over CPUs is maintained. For example, consider 4 (equal weight) tasks and 2 CPUs with a 1:3 distribution, then if F_g_n would simply be 1 (no weight re-distribution) the effective relative weights (W_t') of the tasks in our group would be: CPU0 CPU1 W_g W_g/3 W_g/3 W_g/3 IOW, the lucky task on CPU0 would get an equal amount of weight as all 3 tasks on CPU1 combined. However, with the weight redistribution, this becomes: CPU0 CPU1 W_g/4 W_g/4 W_g/4 W_g/4 All tasks are equal weight (as intended). However, as is already evident from this example, the more CPUs you add, the smaller F_g_n becomes, which creates a disparity against tasks not in our group. Specifically: avg(F_g_n) ~ 1/N This leads to a weight mismatch in the hierarchy. IOW tasks cannot compete fairly across hierarchy levels. *Notably*, what is meant by avg(F_g_n) being proportional to 1/N is that when there are at least N runnable tasks, the average of this fraction tends to 1/N. For a hierarchy of depth d, this gets even worse, since that gets terms on the order of: avg(F_g_n)^d ~ 1/(N^d) Given fixed point arithmetic, this also leads to numerical trouble. However, the meaning of "cpu.weight" is simple and intiutive: the total weight of the cgroup. But as explored above, there is deception in this simplicity. Prepare to add a few alternative methods for distributing weight. Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Link: https://patch.msgid.link/20260605124051.338602724%40infradead.org
2026-06-30sched/fair: Fix overflow in update_tg_cfs_runnable()Chen, Yu C
A divide-by-zero crash is observed when running hackbench: [14697.488452] CPU: 112 UID: 0 PID: 124791 Comm: hackbench Not tainted 7.1.0-rc2+ [14697.492627] RIP: 0010:propagate_entity_load_avg+0x35f/0x3e0 [14697.506799] <TASK> [14697.507411] __dequeue_task+0x2b4/0xc70 [14697.508677] dequeue_task_fair+0x36/0x370 [14697.509047] dequeue_task+0x101/0x2f0 [14697.509426] __schedule+0x1b1/0x1a00 [14697.510868] anon_pipe_read+0x3da/0x450 [14697.511400] vfs_read+0x361/0x390 [14697.512053] __x64_sys_read+0x19/0x30 The divide-by-zero happens here: if (scale_load_down(gcfs_rq->load.weight)) { load_sum = div_u64(gcfs_rq->avg.load_sum, scale_load_down(gcfs_rq->load.weight)); } gcfs_rq->load.weight is an insane large value and is truncated to the lower 32 bits by div_u64, which happen to be 0. Using AI for investigation, the cause is a u32 overflow in update_tg_cfs_runnable(), and flat pickup became a victim when using tg_tasks(): u32 new_sum, divider; ... new_sum = se->avg.runnable_avg * divider; <-- boom The following sequence shows how this triggers the crash: propagate_entity_load_avg() update_tg_cfs_runnable() # u32 overflow corrupts runnable_sum __update_load_avg_cfs_rq() ___update_load_avg() # computes insane runnable_avg update_tg_load_avg() # propagates to tg->runnable_avg update_cfs_group() calc_concur_shares() tg_tasks() # long-to-int truncation, negative nr reweight_entity() # corrupted se->load.weight update_load_add() # corrupted cfs_rq->load.weight propagate_entity_load_avg() update_tg_cfs_load() div_u64() # divide-by-zero Fix by widening new_sum from u32 to u64 (no need to force tg_tasks() to return unsigned long after this fix) Fixes: 95246d1ec80b ("sched/pelt: Relax the sync of runnable_sum with runnable_avg") Assisted-by: Claude:claude-opus-4.6 Signed-off-by: Chen Yu <yu.c.chen@intel.com> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Link: https://patch.msgid.link/a22eea2b-4c4a-4623-9a44-d7b18c0c91c8@intel.com
2026-06-30sched/core: Fix inter-class wakeup_preempt()Peter Zijlstra
The way wakeup_preempt() works since commit 704069649b5b ("sched/core: Rework sched_class::wakeup_preempt() and rq_modified_*()") is that it will call rq->next_class->wakeup_preempt(rq, p) when p is of an equal or higher class, and raise ->next_class when higher. This means that: running idle task wakeup fair-A (next_class == idle) if (sched_class_above(fair, idle)) { wakeup_preempt_idle(fair-A); resched_curr(rq); next_class = fair; } wakeup fair-B (next_class == fair) if (fair == fair) wakeup_preempt_fair(fair-B); (but current is idle) All wakeup_preempt_$class() methods, except for wakeup_preempt_scx() (for whoem this was build) ignore cross-class wakeups by testing if @p is of the right class, but per the above case, it also should check current. This is mostly harmless in the current form, but will lead to trouble with later patches. Fixes: 704069649b5b ("sched/core: Rework sched_class::wakeup_preempt() and rq_modified_*()") Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Link: https://patch.msgid.link/20260626074605.GB2568396%40noisy.programming.kicks-ass.net
2026-06-30fsl/fman: Free init resources on KeyGen failure in fman_init()Haoxiang Li
fman_muram_alloc() allocates initialization resources before initializing the KeyGen block. If keygen_init() fails, the function returns -EINVAL directly and leaves those resources allocated. Free the initialization resources before returning from the KeyGen failure path. Fixes: 7472f4f281d0 ("fsl/fman: enable FMan Keygen") Cc: stable@kernel.org Signed-off-by: Haoxiang Li <haoxiang_li2024@163.com> Reviewed-by: Pavan Chebbi <pavan.chebbi@broadcom.com> Reviewed-by: Breno Leitao <leitao@debian.org> Link: https://patch.msgid.link/20260625004834.3394389-1-haoxiang_li2024@163.com Signed-off-by: Paolo Abeni <pabeni@redhat.com>
2026-06-30mm/mm_init: fix incorrect node_spanned_pagesWei Yang
Current node_spanned_pages is got as a summation of all zone's spanned page in calculate_node_totalpages(). Generally this is good, but if we use kernelcore=mirror, it is would be wrong. Without kernelcore=mirror: The test machine has below memory layout: memory[0x0] [0x0000000000001000-0x000000000009efff], 0x000000000009e000 bytes on node 0 flags: 0x0 memory[0x1] [0x0000000000100000-0x00000000bffdefff], 0x00000000bfedf000 bytes on node 0 flags: 0x0 memory[0x2] [0x0000000100000000-0x00000001bfffffff], 0x00000000c0000000 bytes on node 0 flags: 0x0 And the Zone range is: DMA [mem 0x0000000000001000-0x0000000000ffffff] DMA32 [mem 0x0000000001000000-0x00000000ffffffff] Normal [mem 0x0000000100000000-0x00000001bfffffff] Then we see, with spanned_pages printed: On node 0 spanned_pages: 1835007 totalpages: 1572733 With kernelcore=mirror: The test machine has below memory layout: memory[0x0] [0x0000000000001000-0x000000000009efff], 0x000000000009e000 bytes on node 0 flags: 0x2 memory[0x1] [0x0000000000100000-0x00000000bffdefff], 0x00000000bfedf000 bytes on node 0 flags: 0x2 memory[0x2] [0x0000000100000000-0x000000013fffffff], 0x0000000040000000 bytes on node 0 flags: 0x2 memory[0x3] [0x0000000140000000-0x00000001bfffffff], 0x0000000080000000 bytes on node 0 flags: 0x0 And the Zone range is: DMA [mem 0x0000000000001000-0x0000000000ffffff] DMA32 [mem 0x0000000001000000-0x00000000ffffffff] Normal [mem 0x0000000100000000-0x00000001bfffffff] Device empty Movable zone start for each node Node 0: 0x0000000140000000 Then we see, with spanned_pages printed: On node 0 spanned_pages: 2359295 totalpages: 1572733 The total range of memory on node 0 doesn't change, but the spanned_pages becomes much larger. The reason is when kernelcore=mirror is specified, the range of Zone Normal and Zone Movable would overlap. So the overlapped range would be calculated twice. A wrong node_spanned_pages would effect defer_init(), since each zone_end_pfn is less than pgdat_end_pfn(). As we already passed in node_start_pfn and node_end_pfn, fix this by get it from (node_start_pfn - node_end_pfn) directly. Fixes: 342332e6a925 ("mm/page_alloc.c: introduce kernelcore=mirror option") Signed-off-by: Wei Yang <richard.weiyang@gmail.com> Cc: Yuan Liu <yuan1.liu@intel.com> Link: https://patch.msgid.link/20260622022403.16375-1-richard.weiyang@gmail.com Signed-off-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
2026-06-30USB: serial: keyspan_pda: fix information leakJohan Hovold
The write() callback is supposed to return the number of characters accepted or a negative errno. Since the addition of write fifo support the keyspan_pda implementation will however return the number characters submitted to the device if the write urb is not already in use. If this number is larger than the number of characters passed to write(), the line discipline continues writing data from beyond the tty write buffer. Fix the information leak by making sure that keyspan_pda_write_start() returns zero on success as intended. Fixes: 034e38e8f687 ("USB: serial: keyspan_pda: add write-fifo support") Cc: stable@vger.kernel.org # 5.11 Signed-off-by: Johan Hovold <johan@kernel.org>
2026-06-30mm/memblock: Remove redundant pageblock_align() in free_unused_memmap()Zhen Ni
The assignment `prev_end = pageblock_align(end)` is redundant because `prev_end` was already aligned to pageblock oundaries inside the loop. Since pageblock_align() is a pure function, calling it again with the same input produces the same result. This line was added in commit f921f53e089a ("memblock: align freed memory map on pageblock boundaries with SPARSEMEM"). Remove it to simplify the code. Signed-off-by: Zhen Ni <zhen.ni@easystack.cn> Link: https://patch.msgid.link/20260612031105.3350181-1-zhen.ni@easystack.cn Signed-off-by: Mike Rapoport (Microsoft) <rppt@kernel.org>
2026-06-30Merge drm/drm-next into drm-misc-nextThomas Zimmermann
Backmerging to get drm-misc-next to v7.2-rc1. Signed-off-by: Thomas Zimmermann <tzimmermann@suse.de>
2026-06-30gpio-f7188x: Add support for NCT6126D version BPaul Louvel
The Nuvoton NCT6126D Super-I/O is available in two hardware revisions. According to the manufacturer datasheet revision 2.4, version A reports chip ID 0xD283, while version B reports chip ID 0xD284. The driver currently only recognizes only the version A ID. Version B only contains hardware fixes unrelated to the GPIO functionality, so it can be supported by simply adding its chip ID without any other driver changes. Fixes: 3002b8642f01 ("gpio-f7188x: fix chip name and pin count on Nuvoton chip") Cc: stable@vger.kernel.org Signed-off-by: Paul Louvel <paul.louvel@bootlin.com> Link: https://patch.msgid.link/20260629-gpio-f7188x-nct6126d-version-b-v1-1-a06226c02a2d@bootlin.com Signed-off-by: Bartosz Golaszewski <bartosz.golaszewski@oss.qualcomm.com>
2026-06-30drm/i915/vrr: require valid min/max vfreq for VRRJani Nikula
Ensure the EDID provided min/max vfreq are valid. Most scenarios are already covered (by coincidence) through the checks in intel_vrr_is_capable() and intel_vrr_is_in_range(), but be more explicit about it. At worst, a zero min_vfreq could lead to a division by zero in intel_vrr_compute_vmax(). Discovered using AI-assisted static analysis confirmed by Intel Product Security. Reported-by: Martin Hodo <martin.hodo@intel.com> Fixes: 117cd09ba528 ("drm/i915/display/dp: Compute VRR state in atomic_check") Cc: stable@vger.kernel.org # v5.12+ Cc: Ankit Nautiyal <ankit.k.nautiyal@intel.com> Reviewed-by: Ankit Nautiyal <ankit.k.nautiyal@intel.com> Link: https://patch.msgid.link/20260625131040.1051272-1-jani.nikula@intel.com Signed-off-by: Jani Nikula <jani.nikula@intel.com>
2026-06-30drm/i915/hdcp: require monotonically increasing seq_num_vJani Nikula
The HDCP 2.2 specification requires the seq_num_v to be monotonically increasing, and repeated seq_num_v needs to be treated as an integrity failure. Make it so. For the first message, seq_num_v must be zero, and is already checked. We can only check for less-than-or-equal for the subsequent messages, where hdcp2_encrypted is true. Discovered using AI-assisted static analysis confirmed by Intel Product Security. Reported-by: Martin Hodo <martin.hodo@intel.com> Fixes: d849178e2c9e ("drm/i915: Implement HDCP2.2 repeater authentication") Cc: stable@vger.kernel.org # v5.2+ Cc: Suraj Kandpal <suraj.kandpal@intel.com> Reviewed-by: Suraj Kandpal <suraj.kandpal@intel.com> Link: https://patch.msgid.link/20260625104407.1025614-1-jani.nikula@intel.com Signed-off-by: Jani Nikula <jani.nikula@intel.com> (cherry picked from commit 58a224375c81179b52558c53d8857b93196d2687) Signed-off-by: Joonas Lahtinen <joonas.lahtinen@linux.intel.com>
2026-06-30drm/i915/hdcp: check streams[] bounds before overflowJani Nikula
The data->streams[] overflow check is done after the buffer overflow has already happened. Move the overflow check before the write. Side note, emitting a warning splat with a backtrace might be overkill here, but prefer not changing the behaviour other than not doing the overrun. Discovered using AI-assisted static analysis confirmed by Intel Product Security. Reported-by: Martin Hodo <martin.hodo@intel.com> Fixes: e03187e12cae ("drm/i915/hdcp: MST streams support in hdcp port_data") Cc: stable@vger.kernel.org # v5.12+ Cc: Anshuman Gupta <anshuman.gupta@intel.com> Cc: Suraj Kandpal <suraj.kandpal@intel.com> Reviewed-by: Suraj Kandpal <suraj.kandpal@intel.com> Link: https://patch.msgid.link/20260625170304.1104723-1-jani.nikula@intel.com Signed-off-by: Jani Nikula <jani.nikula@intel.com> (cherry picked from commit 9284ab3b6e776c315883ac2611283d263c9460fd) Signed-off-by: Joonas Lahtinen <joonas.lahtinen@linux.intel.com>
2026-06-30drm/i915: Return NULL on error in active_instanceJoonas Lahtinen
Avoid returning &node->base when node is NULL due to OOM during GFP_ATOMIC allocation. Discovered using AI-assisted static analysis confirmed by Intel Product Security. Reported-by: Martin Hodo <martin.hodo@intel.com> Fixes: bfaae47db3c0 ("drm/i915: make lockdep slightly happier about execbuf.") Cc: Maarten Lankhorst <maarten.lankhorst@linux.intel.com> Cc: Thomas Hellström <thomas.hellstrom@linux.intel.com> Cc: Simona Vetter <simona.vetter@ffwll.ch> Cc: <stable@vger.kernel.org> # v5.13+ Signed-off-by: Joonas Lahtinen <joonas.lahtinen@linux.intel.com> Reviewed-by: Sebastian Brzezinka <sebastian.brzezinka@intel.com> Reviewed-by: Maarten Lankhorst <maarten.lankhorst@linux.intel.com> Link: https://patch.msgid.link/20260624090940.74840-1-joonas.lahtinen@linux.intel.com (cherry picked from commit 6029bc064f0b1bac184203a50fbaaf070fa18832) Signed-off-by: Joonas Lahtinen <joonas.lahtinen@linux.intel.com>
2026-06-30drm/lima: call drm_mm_init() with a valid allocation rangeHenrik Grimler
lima_vm_create() is currently run before va_start and va_end are set up, meaning they are both 0. lima_vm_create() runs drm_mm_init() with them as arguments for the allocator, and if DRM_DEBUG_MM is enabled the DRM_MM_BUG_ON check in drm_mm_init then fires, as seen here on exynos4412-odroid-u2: [ 1.736297] ------------[ cut here ]------------ [ 1.740370] kernel BUG at drivers/gpu/drm/drm_mm.c:931! [ 1.745574] Internal error: Oops - BUG: 0 [#1] SMP ARM [ 1.750697] Modules linked in: [ 1.753734] CPU: 0 UID: 0 PID: 41 Comm: kworker/u16:1 Not tainted 7.0.10-postmarketos-exynos4 #11 PREEMPT [ 1.763372] Hardware name: Samsung Exynos (Flattened Device Tree) [ 1.769446] Workqueue: events_unbound deferred_probe_work_func [ 1.775261] PC is at drm_mm_init+0x9c/0xa4 [ 1.779339] LR is at lima_vm_create+0x144/0x17c [ ... ] Fix the issue by moving the lima_vm_create() call after va_start and va_end are set up. Fixes: a1d2a6339961 ("drm/lima: driver for ARM Mali4xx GPUs") Signed-off-by: Henrik Grimler <henrik.grimler@axis.com> Signed-off-by: Qiang Yu <yuq825@gmail.com> Link: https://patch.msgid.link/20260601-lima-alloc-fix-v1-1-16d3f3b7b780@axis.com
2026-06-30erofs: use more informative s_id for file-backed mountsGao Xiang
For file-backed mounts, set sb->s_id to the MAJOR:MINOR of sb->s_dev (which fstat() will return) so that kernel messages and the sysfs name are more informative rather than just "erofs: (device erofs): ...". Reviewed-by: Hongbo Li <lihongbo22@huawei.com> Signed-off-by: Gao Xiang <hsiangkao@linux.alibaba.com>
2026-06-29Input: maplemouse - fix NULL pointer dereference in open()Florian Fuchs
Commit 555c765b0cc2 ("Input: mouse - drop unnecessary calls to input_set_drvdata") dropped the input_set_drvdata() call in probe because the data appeared to be unused. However, dc_mouse_open() and dc_mouse_close() were using maple_get_drvdata(to_maple_dev(&dev->dev)). This appears to be accessing the data attached to an instance of maple_device structure, while in reality this actually retrieves driver data from the input device's embedded struct device (doing invalid conversion of input device structure to maple device). After input_set_drvdata() was removed, that lookup started returning NULL and opening the input device dereferences mse->mdev. Restore input_set_drvdata() and convert open() and close() to use input_get_drvdata() so the dependency is no longer hidden. Fixes: 6b3480855aad ("maple: input: fix up maple mouse driver") Fixes: 555c765b0cc2 ("Input: mouse - drop unnecessary calls to input_set_drvdata") Signed-off-by: Florian Fuchs <fuchsfl@gmail.com> Link: https://patch.msgid.link/20260628230715.2982552-1-fuchsfl@gmail.com Cc: stable@vger.kernel.org Signed-off-by: Dmitry Torokhov <dmitry.torokhov@gmail.com>
2026-06-30ttm/pool: Use sentinels in debugfsMario Limonciello
The values in the `page_pool_shrink` debugfs file should recognize the sentinels rather than causing an underflow. Closes: https://gitlab.freedesktop.org/drm/amd/-/work_items/5218 Signed-off-by: Mario Limonciello <mario.limonciello@amd.com> Reviewed-by: Tvrtko Ursulin <tvrtko.ursulin@igalia.com> Link: https://patch.msgid.link/20260427162253.682415-1-mario.limonciello@amd.com Signed-off-by: Mario Limonciello <superm1@kernel.org>
2026-06-30netfilter: nftables: restrict checkum update offsetFlorian Westphal
After previous patch, writes to network header are restricted. However, there is another way to manipulate the l3 header: The checksum update function. Restrict this for network header writes, only the ipv4 header is allowed. This needs run-time checks because BRIDGE, INET, NETDEV families can carry l3 headers other than IP. checksum updates to the udp/tcp (l4) headers are not restricted. Signed-off-by: Florian Westphal <fw@strlen.de>
2026-06-30netfilter: nftables: restrict linklayer and network header writesFlorian Westphal
Don't permit arbitrary writes to linklayer and network header data. Several spots in network stack trust header validation performed in ipv4/ipv6 before PRE_ROUTING hook. For linklayer, allow writes for netdev ingress. For other hooks, only allow link layer writes that do not spill into network header. For network header, check the offset/length combinations: - changing dscp requires store at offset 0 for checsum fixups, so make sure ip version + length field isn't altered. - ip6 dscp starts directly after the version field, so make sure it remains 6. Several of these checks could already be done at rule insertion time. Risk is that this might cause ruleset load failures for existing rulesets. With this change such writes are silently skipped and packet passes unchanged. Transport and inner header bases are not checked / restricted. Signed-off-by: Florian Westphal <fw@strlen.de>
2026-06-30netfilter: nfnetlink_queue: restrict writes to network headerFlorian Westphal
nfnetlink_queue doesn't allow selective replacements of some part of the payload, only complete replacement. If the new data is shorter, skb is trimmed, otherwise expanded. Add minimal validation of the new ip/ipv6 header. Check total len matches skb length. Disallow ip option modifications. IPv6 extension headers are also disabled. IP options and exthdrs could be allowed later after validation pass or ip option recompile. Transport header is not checked. Bridge modifications are rejected. Given userspace doesn't even receive L2 headers, use is limited and I don't think there are any users of bridge nfnetlink_queue, let alone users that modifiy payload. Arp isn't supported at all. Signed-off-by: Florian Westphal <fw@strlen.de>
2026-06-30netfilter: nft_fib: reject fib expression on the netdev egress hookTheodor Arsenij Larionov-Trichkine
A fib expression in a netdev egress base chain dereferences nft_in(pkt), NULL on the transmit path, causing a NULL pointer dereference at eval. nft_fib_validate() masks the hook with NF_INET_* values, but netdev hook numbers are a separate enum that aliases them (NF_NETDEV_EGRESS == NF_INET_LOCAL_IN), so an egress chain passes validation and then faults. Add nft_fib_netdev_validate() that limits each result/flag to the netdev hook where the device it reads exists: the input-device cases (OIF, OIFNAME, ADDRTYPE with F_IIF) to ingress, the output-device case (ADDRTYPE with F_OIF) to egress, ADDRTYPE with no device flag to both. Also restrict nft_fib_validate() to NFPROTO_IPV4/IPV6/INET so its NF_INET_* masks are not applied to another family's hooks. Fixes: 42df6e1d221d ("netfilter: Introduce egress hook") Cc: stable@vger.kernel.org Link: https://lore.kernel.org/netfilter-devel/ajxsjcDOnwllMfoR@strlen.de/ Signed-off-by: Theodor Arsenij Larionov-Trichkine <theodorlarionov@gmail.com> Signed-off-by: Florian Westphal <fw@strlen.de>
2026-06-30netfilter: nfnetlink_cthelper: cap to maximum number of expectation per masterPablo Neira Ayuso
If userspace helper policy updates sets maximum number of expectation to zero, cap it to NF_CT_EXPECT_MAX_CNT (255) on updates too. Fixes: 397c8300972f ("netfilter: nf_conntrack_helper: cap maximum number of expectation at helper registration") Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org> Signed-off-by: Florian Westphal <fw@strlen.de>
2026-06-30netfilter: nf_conntrack_sip: validate skb_dst() before accessing itPablo Neira Ayuso
tc ingress and openvswitch do not guarantee routing information to be available. These subsystems use the conntrack helper infrastructure, and the SIP helper relies on the skb_dst() to be present if sip_external_media is set to 1 (which is disabled by default as a module parameter). This effectively disables the sip_external_media toggle for these subsystems without resulting in a crash. Fixes: cae3a2627520 ("openvswitch: Allow attaching helpers to ct action") Fixes: b57dc7c13ea9 ("net/sched: Introduce action ct") Cc: stable@vger.kernel.org Reported-by: Ren Wei <n05ec@lzu.edu.cn> Signed-off-by: Pablo Neira Ayuso <pablo@netfilter.org> Signed-off-by: Florian Westphal <fw@strlen.de>
2026-06-30netfilter: ipset: fix race between dump and ip_set_list resizeXiang Mei
The release path of ip_set_dump_do() and ip_set_dump_done() read inst->ip_set_list via ip_set_ref_netlink(), a plain rcu_dereference_raw() of the array pointer. These run from netlink_recvmsg() without the nfnl mutex and without an RCU read-side critical section. A concurrent ip_set_create() can grow the array: it publishes the new array, calls synchronize_net() and then kvfree()s the old one. Since the dump paths read the array outside any RCU reader, synchronize_net() does not wait for them and the old array can be freed while they still index into it, causing a use-after-free. The dumped set itself stays pinned via set->ref_netlink, so only the array load needs protecting. Take rcu_read_lock() around it, matching ip_set_get_byname() and __ip_set_put_byindex(). BUG: KASAN: slab-use-after-free in ip_set_dump_do (net/netfilter/ipset/ip_set_core.c:1697) Read of size 8 at addr ffff88800b5c4018 by task exploit/150 Call Trace: ... kasan_report (mm/kasan/report.c:595) ip_set_dump_do (net/netfilter/ipset/ip_set_core.c:1697) netlink_dump (net/netlink/af_netlink.c:2325) netlink_recvmsg (net/netlink/af_netlink.c:1976) sock_recvmsg (net/socket.c:1159) __sys_recvfrom (net/socket.c:2315) ... Oops: general protection fault, probably for non-canonical address ... KASAN NOPTI KASAN: maybe wild-memory-access in range [0x02d6...d0-0x02d6...d7] RIP: 0010:ip_set_dump_do (net/netfilter/ipset/ip_set_core.c:1698) Kernel panic - not syncing: Fatal exception Fixes: 8a02bdd50b2e ("netfilter: ipset: Fix calling ip_set() macro at dumping") Cc: stable@vger.kernel.org Reported-by: Weiming Shi <bestswngs@gmail.com> Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Xiang Mei <xmei5@asu.edu> Acked-by: Jozsef Kadlecsik <kadlec@netfilter.org> Signed-off-by: Florian Westphal <fw@strlen.de>
2026-06-30netfilter: nft_set_pipapo: don't leak bad clone into future transactionFlorian Westphal
On memory allocation failure the cloned nft_pipapo_match can enter a bad state: - some fields can have their lookup tables resized while others did not - bits might have been toggled - scratch map can be undersized which also means m->bsize_max can be lower than what is required This means that the next insertion in the same batch can trigger out-of-bounds writes. Furthermore, a failure in the first can result in the bad clone to leak into the next transaction because the abort callback is never executed in this case (the upper layer saw an error and no attempt to allocate a transactional request was made). Record a state for the nft_pipapo_match structure: - NEW (pristine clone) - MOD (modified clone with good state) - ERR (potentially bogus content) Then make it so that deletes and insertions fail when the clone entered ERR state. In case the very first insert attempt results in an error, free the clone right away. Fixes: 3c4287f62044 ("nf_tables: Add set type for arbitrary concatenation of ranges") Cc: stable@vger.kernel.org Reported-and-tested-by: Seesee <cjc000013@gmail.com> Reviewed-by: Stefano Brivio <sbrivio@redhat.com> Signed-off-by: Florian Westphal <fw@strlen.de>
2026-06-30netfilter: nf_conntrack_expect: zero at allocation timeFlorian Westphal
There are occasional LLM hints wrt. leaking uninitialized data to userspace via ctnetlink. Just zero at allocation time, expectations are not frequently used these days. Intentionally keeps _init as-is because we could theoretically support re-init, so add the missing exp->dir there. Signed-off-by: Florian Westphal <fw@strlen.de>
2026-06-29drm/amd/display: use DisplayID panel type in dm_set_panel_typeChenyu Chen
Wire up the newly parsed panel_type from drm_display_info into amdgpu_dm's panel type detection path. When neither the AMD VSDB nor DPCD determines the panel type, fall back to the DisplayID Display Device Technology field to set PANEL_TYPE_LCD or PANEL_TYPE_OLED accordingly. Also expose LCD to userspace via the panel_type connector property. Assisted-by: Copilot:Claude-Opus-4.6 Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com> Reviewed-by: Mario Limonciello (AMD) <superm1@kernel.org> Link: https://patch.msgid.link/20260526030254.1460480-4-chen-yu.chen@amd.com Signed-off-by: Mario Limonciello <superm1@kernel.org>
2026-06-29drm/edid: parse panel type from DisplayID 2.x Display ParametersChenyu Chen
Parse the Display Parameters Data Block (tag 0x21) defined in DisplayID v2.1a Section 4.2.6. Extract the Display Device Technology field from the color depth and device technology byte, which indicates whether the panel uses LCD or OLED technology. Add a panel_type field to struct drm_display_info and populate it during DisplayID iteration so downstream drivers can use it for panel-type-dependent behavior. Add DRM_MODE_PANEL_TYPE_LCD to the UAPI panel type property alongside the existing OLED value. Assisted-by: Copilot:Claude-Opus-4.6 Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com> Reviewed-by: Jani Nikula <jani.nikula@intel.com> Reviewed-by: Mario Limonciello (AMD) <superm1@kernel.org> Link: https://patch.msgid.link/20260526030254.1460480-3-chen-yu.chen@amd.com Signed-off-by: Mario Limonciello <superm1@kernel.org>
2026-06-29drm/edid: extract base section header processing into helperChenyu Chen
Extract the DisplayID base section header logging and non_desktop detection from update_displayid_info() into a dedicated helper, drm_displayid_process_base_section_header(). Remove the break so the iterator walks through all data blocks, preparing for future patches that will parse additional block types within the loop. The helper is called only once for the base section via a base_section_header_processed flag. Since version and primary_use are only captured from the base section, and extension sections carry a primary use of zero per spec, the non_desktop logic is unaffected. No functional change. Assisted-by: Copilot:Claude-Opus-4.6 Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com> Reviewed-by: Jani Nikula <jani.nikula@intel.com> Reviewed-by: Mario Limonciello (AMD) <superm1@kernel.org> Link: https://patch.msgid.link/20260526030254.1460480-2-chen-yu.chen@amd.com Signed-off-by: Mario Limonciello <superm1@kernel.org>
2026-06-29MAINTAINERS: Update Jason Wang's email addressJason Wang
I will use jasowangio@gmail.com for future review and discussion. Signed-off-by: Jason Wang <jasowang@redhat.com> Link: https://patch.msgid.link/20260629014525.16297-1-jasowang@redhat.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-06-29Merge branch 'net-phy-sfp-fix-mii_bus-leak-and-revert-rollball-bridge-probe'Jakub Kicinski
Petr Wozniak says: ==================== net: phy: sfp: fix mii_bus leak and revert RollBall bridge probe v4 tried to fix the RollBall regression from 8fe125892f40 by deferring the bridge probe to PHY discovery time. Maxime Chevallier and Aleksander Bajkowski both tested that on genuine RollBall hardware and confirmed it does not restore PHY detection (the module still is not ready when the probe runs), and the Sashiko static review flagged the same path. So this version drops the deferred-probe patch and instead reverts 8fe125892f40, restoring the pre-regression behaviour for genuine RollBall modules. A proper fix for slow-initializing modules needs per-module init timing (a longer module_t_wait / a per-module quirk) and genuine RollBall hardware to validate; that is better owned as a follow-up by someone with such a module. ==================== Link: https://patch.msgid.link/cover.1782581445.git.petr.wozniak@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-06-29Revert "net: phy: sfp: probe for RollBall I2C-to-MDIO bridge in mdio-i2c"Petr Wozniak
This reverts commit 8fe125892f40 ("net: phy: sfp: probe for RollBall I2C-to-MDIO bridge in mdio-i2c"). That commit added a RollBall bridge probe at MDIO bus creation time, in i2c_mii_init_rollball(), to avoid a multi-minute PHY probe retry loop on modules without a bridge (e.g. RTL8261BE). The probe runs in SFP_S_INIT, before genuine RollBall modules have finished their firmware/bridge initialization, so the bridge does not yet answer CMD_READ/CMD_DONE. The probe times out, mdio_protocol is set to MDIO_I2C_NONE, and PHY detection is then skipped for genuine RollBall modules that worked before the commit. This was confirmed on hardware by Maxime Chevallier and Aleksander Bajkowski: their RollBall modules no longer detect a PHY, and work again on v7.0 (before the bridge probing was introduced). The Sashiko static review flagged the same path. Deferring the probe to PHY discovery time does not fix it either: at that point a slow module may still be initializing, so the probe still returns -ENODEV. A proper fix needs per-module init timing (a longer module_t_wait or a per-module quirk, per SFF-8472 the host must also wait at least 300 ms after insertion), which requires genuine RollBall hardware to develop and validate. Revert to restore the previous, working behaviour in the meantime. The RTL8261BE retry-loop latency that the reverted commit addressed is handled in our downstream tree, so reverting upstream is safe on our side. Fixes: 8fe125892f40 ("net: phy: sfp: probe for RollBall I2C-to-MDIO bridge in mdio-i2c") Reported-by: Aleksander Bajkowski <olek2@wp.pl> Suggested-by: Maxime Chevallier <maxime.chevallier@bootlin.com> Link: https://lore.kernel.org/netdev/20260624084814.20972-1-petr.wozniak@gmail.com/ Signed-off-by: Petr Wozniak <petr.wozniak@gmail.com> Tested-by: Maxime Chevallier <maxime.chevallier@bootlin.com> Reviewed-by: Maxime Chevallier <maxime.chevallier@bootlin.com> Link: https://patch.msgid.link/23e3931915c3ed2a14cec95f1490e43d30b225e8.1782581445.git.petr.wozniak@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org>