arXiv Query: search_query=all:"software testing"&id_list=&start=0&max_results=10

フィード

記事のアイキャッチ画像
MiniRep: Robust Reputation-Based Aggregation for Multi-Agent Debate
arXiv Query: search_query=all:"software testing"&id_list=&start=0&max_results=10
Autonomous agents powered by large language models (LLMs) are rapidly evolving into an open agentic ecosystem. To support trustworthy collaboration, industry initiatives increasingly assess agent reputation from past behavior and provide performance leaderboards. However, reputation derived from past performance may not reliably predict an agent's behavior on new tasks, particularly when malicious agents can adapt their behavior and influence other agents during collaboration. We study reputation in multi-agent debate (MAD), where multiple agents answer the same query, debate to improve their answers, and aggregate them into a final output. We present MiniRep, a reputation-based aggregation system for MAD under malicious agents. To ground our threat model in established research, we construct an attack taxonomy drawing on reputation-system attacks and software-testing mutation operators, covering strategic exploitation of reputation and subtle corruption of agent proposals. Guided by t
5日前
記事のアイキャッチ画像
Exploring Emotional Intelligence in Software Testing
arXiv Query: search_query=all:"software testing"&id_list=&start=0&max_results=10
Background: Emotional Intelligence (EI) is the ability to recognise, understand, and manage one's own and others' emotions. Software testers deliver judgements about colleagues' work under deadlines they do not control, and prior work on emotion in software engineering has mostly studied developers. Aims: To explore how software testers describe the part EI plays in their day-to-day work, in communication and conflict within the team, and in responding to requirements volatility. Method: Semi-structured interviews with 16 software testers in Sweden working in teams that use agile practices, across aviation, automotive, healthcare, IT services, administration, banking and pharmaceuticals, analysed with reflexive thematic analysis informed by Goleman's EI framework. Results: Three themes. Testers described regulating stress under deadline pressure and drawing motivation from recognition, clarity and autonomy; managing the daily delivery of critical findings to colleagues so that trust su
6日前
記事のアイキャッチ画像
Mubric: Mutation Testing-Guided Rubric Generation for LLM Evaluation
arXiv Query: search_query=all:"software testing"&id_list=&start=0&max_results=10
Rubric-based evaluation is widely used to assess LLM-based systems by decomposing response quality into task-specific scoring criteria. However, automatically generating rubrics that reliably capture task-specific quality requirements remains challenging. We introduce Mubric, a mutation testing-guided approach to rubric generation. Mutation testing, a classic software testing methodology, evaluates a test suite by injecting faults into programs and checking whether the tests detect them. We draw an analogy between test suites and rubrics: if a rubric captures an important quality requirement, introducing a corresponding defect into an otherwise high-quality response should reduce its score. Mubric first mines common defects from real pairs of preferred and dispreferred responses and abstracts these defects into reusable mutation operators, each specifying how to introduce a particular type of response defect. For a new task, it applies relevant operators to a reference response, checks
6日前
記事のアイキャッチ画像
What Makes In-Context Examples Effective for Code Generation?
arXiv Query: search_query=all:"software testing"&id_list=&start=0&max_results=10
In-Context Learning (ICL) has emerged as a promising solution to enhance the code generation capabilities of Large Language Models (LLMs) by incorporating code examples inside the prompt to let LLMs learn from demonstrations. However, despite their effectiveness gains, it remains unclear which specific properties of ICL-provided code examples (e.g., solution insight, essential contextual information, identifier naming styles, code formatting) drive these gains. This paper systematically investigates the impact of different sources and internal features of code examples on ICL for code generation through controlled experiments on contest-style programming questions and repository-level tasks. Our results show that while LLMs struggle to extract generalizable problem-solving insights from provided solutions to similar questions or repository snippets, their retrieval-augmented ICL performance can significantly benefit from explicit contextual information, such as input/output demonstrati
1年前
記事のアイキャッチ画像
What Output-Equivalence Oracles Miss: An Empirical Study of Equivalence-Invisible Bug Fixes in Quantum Transpilers
arXiv Query: search_query=all:"software testing"&id_list=&start=0&max_results=10
Quantum compilers face the test oracle problem, judged by an output-equivalence oracle: the compiled circuit must compute the same unitary as the original, modulo global phase and a qubit-layout permutation. This oracle, by construction, checks only that semantic map, not the circuit's own layout, permutation, or phase records: a defect there, or in a fixed-seed run's determinism, can pass unseen though the record is public. This empirical software engineering study of quantum transpiler correctness uses repository mining to measure how often this happens in real merged compiler fixes: a systematically identified corpus of Qiskit transpiler bug-fixes, classified by an independently dual-coded, source-validated fault-manifestation taxonomy. Nineteen of 68 fixes (28%, 95% Wilson CI 19-40%) repair faults invisible to this equivalence screen, even one augmented with compilation-validity, circuit-quality, and performance checks, and an extended 104-fix corpus over a wider window holds at th
23日前
記事のアイキャッチ画像
WebCraftBench: Evaluating Web Application Generation from a Software Testing Perspective
arXiv Query: search_query=all:"software testing"&id_list=&start=0&max_results=10
Human evaluation provides a direct measure of the quality of LLM-generated web applications. However, fitting human judgments through automated evaluation remains challenging. Static benchmarks can credit functionality that exists in source code but is unreachable at runtime. Interactive benchmarks exercise the application, yet incomplete exploration can cause them to miss implemented functionality and confound application defects with agent execution failures. To address these limitations, we propose WebCraftBench, an interactive benchmark for evaluating web application generation from a software testing perspective. WebCraftBench instruments each generated application and uses code coverage to guide an agent in exploring its functionality through user-simulated interactions. It then abstracts the interaction trace into a state-transition graph and evaluates the application along three dimensions: visual aesthetics, usability, and requirement alignment. By separating exploration from
21日前
記事のアイキャッチ画像
Search-Based Metamorphic Testing of Vision-Language Models in Autonomous Underwater Robotic Software
arXiv Query: search_query=all:"software testing"&id_list=&start=0&max_results=10
Our industry partner focuses on quality assurance for industrial systems across multiple domains, including maritime systems, such as overwater vessels and autonomous underwater robots (AURs). Despite the strong performance of vision-language models (VLMs) in scene understanding, image captioning, and object recognition, their use in AUR software operating in underwater environments is underexplored. Therefore, in this context, it is important to evaluate the quality of VLMs for integration into AUR software and, so, automated software testing tools are needed to assess their suitability and improve their dependability. To this end, we propose a search-based metamorphic testing approach (MetaVLM) that identifies a minimal set of transformations on underwater images to induce incorrect model predictions, thereby revealing VLM failures. We employ NSGA-II as a multi-objective search algorithm and evaluate it over open-source VLMs, BLIP and CLIP, against a random search baseline. Results d
20日前
記事のアイキャッチ画像
Is Self-Admitted Technical Debt Tested? An Empirical Study of Coverage, Co-change, and Impact
arXiv Query: search_query=all:"software testing"&id_list=&start=0&max_results=10
When developers write a TODO or FIXME comment, they are explicitly admitting that the code is suboptimal: a built-in warning that this logic deserves extra scrutiny. Yet it is an open question whether Self-Admitted Technical Debt (SATD) actually receives that scrutiny in the form of software testing. We aim to characterize the relationship between SATD and testing across three dimensions: the extent to which SATD-affected code is covered by existing tests, whether developers synchronize test additions with debt resolution, and whether such testing affects the long-term observability of resulting defects. For that, we conducted an empirical study on eight open-source Java projects, analyzing test coverage of 784 SATD instances identified in the latest releases and performing a longitudinal examination of 5,175 SATD removal events. Our results show that while 60.7% of SATD-affected code is covered by existing test suites, developers rarely synchronize test modifications with debt resolut
24日前
記事のアイキャッチ画像
Deep Learning-Based Detection of Electrical Faults and Power Quality Disturbances in Aerospace Power Systems
arXiv Query: search_query=all:"software testing"&id_list=&start=0&max_results=10
More Electric Aircraft require fast and reliable monitoring of high-frequency electrical networks, yet most power quality disturbance and fault diagnosis methods are developed for conventional 50 or 60 Hz grids. This work presents a hardware-aware deep learning framework for multiclass detection of electrical faults and power quality disturbances in a 400 Hz aerospace power system. A high-fidelity simulation model inspired by the Boeing 787 electrical architecture generates voltage and current waveforms for 21 normal, disturbance, switching, open-circuit, and short-circuit conditions. Two datasets, each containing 73,500 samples, are formed from one-dimensional time-series signals and short-time Fourier transform time-frequency representations. Signal-processing augmentation, domain randomization, and class-specific generative adversarial networks increase waveform diversity, and the time-series dataset is released through IEEE DataPort. We compare 1D and 2D convolutional neural networ
1ヶ月前
記事のアイキャッチ画像
How effective are traditional test criteria at detecting bugs in large language models generated code?
arXiv Query: search_query=all:"software testing"&id_list=&start=0&max_results=10
Test adequacy criteria are widely used to evaluate and guide software testing. Although prior research has extensively examined these criteria using human-written programs, faults, and tests, the increasing adoption of Large Language Models (LLMs) for code generation raises important questions about their effectiveness in detecting LLM-induced faults. To investigate this, we conduct an empirical study involving 5 LLMs and 4 benchmarks, simulating end-to-end workflows in which both code and tests are automatically generated. We collect 6,000+ faulty program instances and evaluate the effectiveness and efficiency of 3 widely used adequacy criteria: statement coverage, branch coverage, and mutation testing. Our findings reveal several key insights. First, most faults introduced by LLMs are relatively trivial to catch. Second, the challenging faults are difficult to trigger using either traditional coverage-based or mutation-based criteria. Third, actual fault detection rates remain extrem
1ヶ月前