[SPIRV] Infer external pointer parameter types from mangled names (#230342)
Pointer parameters of external declarations can fall back to i8 because
there is no body for pointee type inference.
Recover scalar and vector pointees from unambiguous mangled signatures
and use them consistently for function types
and calls.
Preserve explicit pointee attributes and existing builtin handling. Keep
the byte-pointer fallback for unsupported types, vector ranks, and
signatures whose parameters cannot be safely matched to LLVM arguments.
[RISCV] Use ComplexPattern for matching addi+brind. (#229594)
This only requires 1 isel pattern per BRIND type and we're able
to capture the add vs ptradd difference between SelectionDAG and
GISel in the C++ code.
Assisted-by: Claude
[Clang] [NFC] Add `APValue::visit()` (#230243)
This patch adds a helper function to `APValue` that walks the value
recursively and visits all nested values. This functionality is
currently used by `LinkageComputer::getLVForValue()` and it looks like I
will be needing it for a patch soon.
Assisted-by: Codex
[Offload][omp] Get Plugin active count and Device UID through liboffload (#230109)
Reimplement PluginManager::getNumActivePlugins and Device getDeviceUid
through liboffload.
Assisted by Claude.
[SLP]Drop nneg when MinBW changes the cast operand type
MinBW may narrow the operand of a vectorized cast, so the nneg flag
propagated from the original scalar cast no longer holds for the new
operand type.
Fixes #230443
Reviewers:
Pull Request: https://github.com/llvm/llvm-project/pull/230671
Allow quota to be changed even when over quota
ZFS allows datasets to use slightly more space than their quota,
typically one write operation totaling less than 1MB. Once in this state
of using more space than the quota, subsequent writes will fail with
ENOSPC. The quota can also be changed to more than the space used (with
`zfs set quota=...`). The quota can not be tightened (decreased) such
that the dataset is in an over-quota state, which is a design decision
dating back to before version 1.
The problem is that when in the state of using more space than the
quota, the quota can not be set to its current value, or relaxed
(increased) to a value that is less than the space used. Since these
operations don't make anything worse in terms of being over-quota, they
should be allowed. This enables `zfs set quota=` to be idempotent.
Original-patch-by: Matthew Ahrens <matt at mahrens.org>
External-issue: https://www.illumos.org/issues/18300
External-issue: https://www.illumos.org/issues/18332
[6 lines not shown]
[Offload][Omp] Ensure errors are consumed when debug is off (#230668)
When debug is off (either not enabled at compile or runtime) the code
consuming the Error (e.g., toString) is not executed and the program can
potentially abort with an unconsumed error (reported by @ro-i).
Claude scanned for ocurrences of this pattern and only the messages
introduced by #226450 seem to be affected.
Note that messages using `REPORT()` unlike `ODBG` are safe.
[Clang][Driver] Use offload LTO mode when embedding device bitcode (#230644)
With `-fno-offload-lto`, `-fembed-bitcode=all` collapsed the device
compile into a backend job that drops HIP includes
Regressed is introduced #201155 (and later in #202736)
graphics/nvidia-drm-kmod: Drop obsolete patch
On commit ports 3bcc752b87a2, supports for 15.0-RELEASE are dropped
in all ports.
Afterwards, graphics/nvidia-drm-kmod/Makefile.common still has
a conditional that could attempt to apply deleted patch.
Remove the whole affected conditional to avoid potential breakage
on attempting to override version.
PR: 299111
Differential Revision: https://reviews.freebsd.org/D60289
AMDGPU: Index loads by workitem id in tests shared with r600
These tests relied on -amdgpu-scalarize-global-loads=false to select
vector loads from uniform pointer arguments. They also have r600 run
lines, so keep the kernels and index the input pointers by the workitem
id instead. Also fix shl_v2i16 not using its computed pointers, and
v_shl_32_i64 using the workgroup id. Also add some uniform variants
of some cases.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
Linux: Sleep instead of spinning when zfs_zget() races eviction
When igrab() fails because the VFS is evicting the inode, zfs_zget()
drops its locks, calls cond_resched() and retries. cond_resched()
only yields when a reschedule is already pending, so every lookup of
an inode that is waiting in a long dispose_list() spins at full speed,
taking and dropping the znode hold locks and allocating a znode_hold_t
on each pass. With enough of them the thread doing the eviction is
starved of both CPU and those locks, and the eviction stalls. On a
24-thread NAS serving ~180 rsync clients this turned into a livelock
with zero pool I/O that needed a power cycle.
Sleep for one tick before retrying instead, as xfs_iget() does when it
races the same VFS teardown. The retry loop is otherwise unchanged and
no locks are held while sleeping.
Co-Authored-By: Claude Opus 5.5 <noreply at anthropic.com>
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Karl Schulze <karl at taniustech.com>
[2 lines not shown]
Linux: Feed the superblock shrinker in batches from zfs_prune()
zfs_prune() hands the whole ARC prune request to a single
super_cache_scan() call. When arc_evict() asks for hundreds of
thousands or millions of objects, prune_icache_sb() isolates up to
that many unused inodes, marks every one of them I_FREEING, and only
then evicts them one at a time from dispose_list(). For as long as it
takes to work through that list, a lookup of any of those inodes fails
igrab() in zfs_zget() and has to retry.
The kernel's own reclaim never builds such a list: do_shrink_slab()
calls scan_objects() in chunks of shrinker->batch objects (1024 for
superblocks). Do the same in zfs_prune(), so that at most one batch
of inodes is I_FREEING at a time. The total number of objects scanned
per prune request is unchanged.
Co-Authored-By: Claude Opus 5.5 <noreply at anthropic.com>
Reviewed-by: Brian Behlendorf <behlendorf1 at llnl.gov>
Signed-off-by: Karl Schulze <karl at taniustech.com>
[2 lines not shown]
AMDGPU: Stop relying on -amdgpu-scalarize-global-loads=false in ALU tests (#229924)
Convert kernels which loaded operands from pointer arguments into functions taking the
operands as arguments. Scalar kernel arguments become inreg arguments. Where a kernel
is still useful, index the loads by the workitem id so they remain vector loads. This also fixes
a few tests where the workitem id index was computed but unused, or where the uniform
workgroup id was used for indexing in VALU tests.
Co-authored-by: Claude Opus 5.5 <noreply at anthropic.com>
[SLP]Fix gather of extractelements from different-sized vectors (#230664)
Check the extract index against its own vector and exit early if the
size does not match the selected operands.
Fixes #230442
CI: verify "failing test, then fix" PRs
A bug fix is easiest to review and keep fixed when the PR first adds a
test that shows the bug, then fixes it. Add a workflow that checks
this pattern: for each commit that only changes tests/ and is followed
by a commit changing code, build that commit in a QEMU VM and check its
new or changed tests fail (or crash or hang the kernel), then build the
PR head and check the same tests pass without kernel errors.
The workflow reuses the qemu-* scripts to set up, build and boot one
test VM per build, runs each test on its own with zfs-tests.sh -t, and
restarts the VM after a crash. PRs without a test commit finish after
the detect job.
The detector gives each job 15 minutes plus 30 per test (at most 240),
and the per-test watchdog in failfirst-tests.sh drops from 40 to 25
minutes: zfs-tests.sh -t stops a test after 600 seconds, and the rest
covers the group's setup and cleanup. A job with one test times out
after 45 minutes.
[3 lines not shown]
RegisterPressure: Detect dead physreg defs from LiveIntervals
This reverts the remainder of #222627, which was partially reverted by
on dead flags. This is a prerequisite to deleting LiveVariables.
When constructing the PressureDiff for an instruction during scheduling
DAG construction, dead defs were only recognized from the dead flag on the
operand. This implicitly relied on preprocessing done by LiveVariables to
fixup inconsistent dead flags with overlapping registers in other operands.
Dead flags have no verifier-enforced rules and are thus unreliable.
Before LiveVariables, consider this example:
dead $eax = MOV32r0 implicit-def dead $eflags, implicit-def $rax
; $rax is never used
$rax is never used, but only the $eax def is dead-flagged and the overlapping
implicit-def $rax is not. The shared $eax register units are covered by the
non-dead $rax def and so are counted as live defs. That shared unit is then
[26 lines not shown]
[clang] Fix crash on unterminated __identifier (#229156)
With Microsoft extensions enabled, e.g. via `-fms-compatibility`, the
following input triggers an assertion:
```cpp
# 1 __identifier(foo
```
The `__identifier` handler consumes the end-of-directive token, `eod`,
while looking for `)`, but still returns `foo` to the line-marker
handler. The latter rejects `foo` as an invalid filename and tries to
discard the remaining tokens up to the end of the directive. Since `eod`
has already been consumed, it encounters `eof` and triggers an
assertion.
Return `eod`, `eof`, and `annotation` tokens to the caller when
recovering from a missing `)`.
Fixes #222310
textproc/py-hieroglyph: Remove expired port
2026-10-07 textproc/py-hieroglyph: Broken with textproc/py-sphinx 7 and newer, last upstream release in 2020
[AMDGPU] Add async and tensor waits to isWaitcnt
S_WAIT_ASYNCCNT and S_WAIT_TENSORCNT were missing, so passes treated
them as ordinary instructions. S_SET_VGPR_MSB was not coissued past
them, and SIPreEmitPeephole could drop the execz skip around them.
Drop the now redundant opcode checks in classifyFlavor.
Change-Id: If34ce4fa7d3c3f86b74015b4b355c444e224d6dc
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
[LV] Add tests for FindLast reductions with nested phis (NFC). (#230640)
Add tests where the data value of a FindLast reduction is itself a phi
that yields the reduction phi on some paths, e.g. via a nested if. Both
are currently miscompiled, as the outer condition is used as the
condition for updating the reduction.
Also enable masked memory ops and regenerate checks with the latest UTC
version to make checks more compact without loss of generality.