[lldb] Fix deadlock when a REPL expression hits a breakpoint (#227400)
When a REPL expression stops at a breakpoint,
REPL::IOHandlerInputComplete calls RunIOHandlerAsync to drop into the
command interpreter while it still holds the error stream lock.
RunIOHandlerAsync then takes the IOHandler stack mutex. At the same
time, the event handler thread reports the stop through
IOHandlerStack::PrintAsync, which takes the IOHandler stack mutex first
and the output mutex second. The opposite lock order can deadlock,
leaving lldb hung after it prints "Execution stopped at breakpoint."
The lock scope that covers the call was introduced in 5007dd9d0945
(#183600). Keep printing under the lock, but defer RunIOHandlerAsync
until the lock is released.
This patch adds a test that exercises this path with the C REPL. Because
the deadlock is a race, the test only catches a regression
intermittently: without the fix it hung in about 6% of runs under
parallel load, and in every run when a delay was injected inside the
[2 lines not shown]
[obj2yaml] Suppress leak check on abnormal exit (#227444)
This is a common issue with lsan after exit().
exit() is noreturn, so we cannot expect that the compiler will
preserve pointers to allocations done by callers.
Fixes https://lab.llvm.org/buildbot/#/builders/169/builds/27022
Assisted-by: Gemini
TwoAddressInstructions: Fix subranges straddling the INSERT_SUBREG subreg index (#227427)
Rewriting INSERT_SUBREG into a subregister COPY narrows the def from the
whole register down to one subreg index. The live interval fixup for that only
handled subranges entirely disjoint from the inserted lane mask, so a partially
overlapping subrange keeps a value whose def no longer writes all of its lanes.
The stale value makes the lanes outside the inserted subregister look
defined by the COPY, so the copy that actually provides them ends up
dead and its value is lost. Refine the subranges against the subreg index
before narrowing the def, leaving each one either fully redefined by the COPY
or untouched by it.
Fixes a miscompile reported on RISC-V and restores the pre-6286f77214be
output of Thumb2/mve-vst2.ll, Thumb2/mve-vst3.ll and PowerPC/dmr-enable.ll.
Co-authored-by: Claude Opus 5 <noreply at anthropic.com>
[CIR] Coerce record returns from memory like CreateCoercedLoad
Read a coerced record return straight from the alloca it was loaded
from. If that isn't possible, copy the record's bytes into a coercion
slot and read from there. Neither path stores the record as a value.
A union stored as a value keeps only its storage member and loses the
bytes that are padding in that member.
Assisted-by: Claude Code / Claude Opus 5.5
Use new-style triples in SRAMECC mode test
Resolve the processor from the triple subarch as requested in review.
Change-Id: Ibbfd674a974fa19e71ef08497c27010e7a452c43
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
bhyve: rtc_pl031: Fix PeriphID and CellID values
PeriphID and CellID values are determined by macros which take an
index. They currently receive a bus offset which has a stride of 4 bytes.
This causes the ID1-3 registers to report incorrect values.
Scale the offset before passing it to the macro to fix this.
Tested with kvm-unit-tests/arm/pl031.
Signed-off-by: Kajetan Puchalski <kajetan.puchalski at arm.com>
Reviewed by: jrtc27
Fixes: 014d7082a239 ("bhyve: Implement a PL031 RTC on arm64")
MFC after: 1 week
Pull Request: https://github.com/freebsd/freebsd-src/pull/2358
Closes: https://github.com/freebsd/freebsd-src/pull/2358
(cherry picked from commit a554906ea44c26925730a25263e64890d48d2b36)
[HLSL][Matrix] Convert loaded bool matrices to `i1` (#226599)
Fixes #226308.
This PR converts loaded bool matrices to their `i1` representation
before use, like bool vectors do, and updates the affected matrix
codegen tests.
It also splits matrix tests out of `or.hlsl` into its own `or_mat.hlsl`
file to match the convention of the other matrix intrinsic tests.
Assisted-by: Claude Opus 4.8
[AMDGPU] Validate scale_sel in v_cvt_scale_*
These instructions can be block16 or block32 depending on the target
and scale_sel bits. Block16 is not supported in strict mode.
Re-enable the rest of the instructions in the strict mode but validate
the scale selector.
[AMDGPU][GlobalISel] Keep typed LLTs in RegBankLegalize combines
The S1 cleanup combines in AMDGPURegBankLegalize built new registers
with an untyped s32, missed when lowering switched to extended LLTs.
The regbank combiner later unified these with typed registers, so
constrainRegAttrs retyped e.g. an i32 G_CTPOP result to s32, which
failed instruction selection.
Take the type from the source or destination register instead.
Change-Id: I06cfb7278af045dbaa3ca0cc6c77a3938f4cef01
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
[lldb] Bounds-check ELF relocation offsets (#227046)
Relocations applied to ELF debug sections were written at r_offset
without checking that the target fits in the section. A malformed object
file could therefore write anywhere relative to the file buffer.
Check every relocation target against the section data. Skip and report
relocations that fall outside it.
rdar://186921964
[CIR] Lower variadic arguments through an indirect call
On x86_64, CallConvLowering now classifies an indirect call with
ellipsis arguments from its own operands, as it does a direct variadic
call. Those arguments stay out of the retyped callee's function type, as
in classic CodeGen.
The cir.call verifier now checks indirect calls against the callee
pointer's function type, and CIRGen's asm-label redirect keeps a
variadic declaration's ellipsis instead of dropping it.
Assisted-by: Claude Code / claude-opus-5-5
[AMDGPU][GlobalISel] Keep typed LLTs in RegBankLegalize combines
The S1 cleanup combines in AMDGPURegBankLegalize built new registers
with an untyped s32, missed when lowering switched to extended LLTs.
The regbank combiner later unified these with typed registers, so
constrainRegAttrs retyped e.g. an i32 G_CTPOP result to s32, which
failed instruction selection.
Take the type from the source or destination register instead.
Change-Id: I06cfb7278af045dbaa3ca0cc6c77a3938f4cef01
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply at anthropic.com>
[AMDGPU] Validate scale_sel in v_cvt_scale_*
These instructions can be block16 or block32 depending on the target
and scale_sel bits. Block16 is not supported in strict mode.
Re-enable the rest of the instructions in the strict mode but validate
the scale selector.
[AMDGPU][GlobalISel] Fold neg/abs modifiers when mad-mix selects the low half
selectVOP3PMadMixModsImpl re-runs the fneg/fabs match after rewriting Src
to the 32-bit register the 16-bit value is a half of, but only did so on
the isExtractHiElt path, not for isExtractLoElt.
The two halves are not symmetric. An fneg/fabs of a 32-bit float only
touches bit 31, which is the sign bit of the high half, so folding it
into a modifier on the selected high half is correct. The low half's sign
bit is bit 15, which such an fneg/fabs leaves alone, so only a modifier
that acts on each 16-bit element can be folded there. So the source must be
a 2 x 16-bit vector to fold it.
[InstCombine] Handle select-like i1 sext/zexts in canonicalizeClampLike (#227067)
canonicalizeClampLike matches a select of a select. But if the inner
select has an i1 result type it will be canonicalized to a sext or zext
of the condition. In FFmpeg there is a clamping pattern that shifts the
inverted bits instead of the negated, and the inner select ends up being
canonicalized this way: https://godbolt.org/z/5q8bx69x3
This teaches canonicalizeClampLike to use m_SelectLike to catch these
cases too.