亚洲黄色网站不卡免费_一区二区三区精品国产亚洲_亚洲国产极品专区在线观看_老太婆性杂交毛片_精品99一区二区三区麻豆_成年人视频在线观看免费_亚洲AV成人一区二区三区高清

This paper introduces innovative benchmarks to evaluate Vision-Language Models (VLMs) in real-world zero-shot recognition tasks, focusing on the granularity and specificity of prompting text. We propose a unique evaluation protocol using adapted ImageNet and MS-COCO datasets to assess models' consistency in recognizing concepts at varying granularity levels and their sensitivity to the specificity of language inputs. Our extensive evaluation reveals that state-of-the-art VLMs, including contrastive models like CLIP, struggle with granularity and are sensitive to text specificity, impacting their effectiveness in open-world settings. This comprehensive study, a first in evaluating VLMs from these perspectives, provides valuable insights and tools for the community, highlighting the limitations and paving the way for enhanced models with better generalization in zero-shot recognition.

相關內容

MoDELS

關注 43

ACM/IEEE第23屆模型驅動工程語言和系統國際會議，是模型驅動軟件和系統工程的首要會議系列，由ACM-SIGSOFT和IEEE-TCSE支持組織。自1998年以來，模型涵蓋了建模的各個方面，從語言和方法到工具和應用程序。模特的參加者來自不同的背景，包括研究人員、學者、工程師和工業專業人士。MODELS 2019是一個論壇，參與者可以圍繞建模和模型驅動的軟件和系統交流前沿研究成果和創新實踐經驗。今年的版本將為建模社區提供進一步推進建模基礎的機會，并在網絡物理系統、嵌入式系統、社會技術系統、云計算、大數據、機器學習、安全、開源等新興領域提出建模的創新應用以及可持續性。官網鏈接： · Extensibility · Performer · 可辨認的 · Pivotal（公司） ·

2024 年 3 月 11 日

Practical Implementation of RIS-Aided Spectrum Sensing: A Deep Learning-Based Solution

Sefa Kayraklik,Ibrahim Yildirim,Ertugrul Basar,Ibrahim Hokelek,Ali Gorcin

from arxiv, Accepted in IEEE Systems Journal, Copyright IEEE

This paper presents reconfigurable intelligent surface (RIS)-aided deep learning (DL)-based spectrum sensing for next-generation cognitive radios. To that end, the secondary user (SU) monitors the primary transmitter (PT) signal, where the RIS plays a pivotal role in increasing the strength of the PT signal at the SU. The spectrograms of the synthesized dataset, including the 4G LTE and 5G NR signals, are mapped to images utilized for training the state-of-art object detection approaches, namely Detectron2 and YOLOv7. By conducting extensive experiments using a real RIS prototype, we demonstrate that the RIS can consistently and significantly improve the performance of the DL detectors to identify the PT signal type along with its time and frequency utilization. This study also paves the way for optimizing spectrum utilization through RIS-assisted CR application in next-generation wireless communication systems.

MoDELS · INTERACT · Performer · Pivotal（公司） · state-of-the-art ·

2024 年 3 月 11 日

EMO-SUPERB: An In-depth Look at Speech Emotion Recognition

Haibin Wu,Huang-Cheng Chou,Kai-Wei Chang,Lucas Goncalves,Jiawei Du,Jyh-Shing Roger Jang,Chi-Chun Lee,Hung-Yi Lee

from arxiv, webpage: //emosuperb.github.io/

Speech emotion recognition (SER) is a pivotal technology for human-computer interaction systems. However, 80.77% of SER papers yield results that cannot be reproduced. We develop EMO-SUPERB, short for EMOtion Speech Universal PERformance Benchmark, which aims to enhance open-source initiatives for SER. EMO-SUPERB includes a user-friendly codebase to leverage 15 state-of-the-art speech self-supervised learning models (SSLMs) for exhaustive evaluation across six open-source SER datasets. EMO-SUPERB streamlines result sharing via an online leaderboard, fostering collaboration within a community-driven benchmark and thereby enhancing the development of SER. On average, 2.58% of annotations are annotated using natural language. SER relies on classification models and is unable to process natural languages, leading to the discarding of these valuable annotations. We prompt ChatGPT to mimic annotators, comprehend natural language annotations, and subsequently re-label the data. By utilizing labels generated by ChatGPT, we consistently achieve an average relative gain of 3.08% across all settings.

Analysis · 估計/估計量 · 可辨認的 · 值域 · LIDAR ·

2024 年 3 月 8 日

Prepared for the Worst: A Learning-Based Adversarial Attack for Resilience Analysis of the ICP Algorithm

Ziyu Zhang,Johann Laconte,Daniil Lisus,Timothy D. Barfoot

from arxiv, 8 pages (7 content, 1 reference). 5 figures, submitted to the IEEE Robotics and Automation Letters (RA-L)

This paper presents a novel method to assess the resilience of the Iterative Closest Point (ICP) algorithm via deep-learning-based attacks on lidar point clouds. For safety-critical applications such as autonomous navigation, ensuring the resilience of algorithms prior to deployments is of utmost importance. The ICP algorithm has become the standard for lidar-based localization. However, the pose estimate it produces can be greatly affected by corruption in the measurements. Corruption can arise from a variety of scenarios such as occlusions, adverse weather, or mechanical issues in the sensor. Unfortunately, the complex and iterative nature of ICP makes assessing its resilience to corruption challenging. While there have been efforts to create challenging datasets and develop simulations to evaluate the resilience of ICP empirically, our method focuses on finding the maximum possible ICP pose error using perturbation-based adversarial attacks. The proposed attack induces significant pose errors on ICP and outperforms baselines more than 88% of the time across a wide range of scenarios. As an example application, we demonstrate that our attack can be used to identify areas on a map where ICP is particularly vulnerable to corruption in the measurements.

Integration · Principle · Continuity · 簇 · Kubernetes ·

2024 年 3 月 8 日

vSPACE: Voting in a Scalable, Privacy-Aware and Confidential Election

Se Elnour,William J Buchanan,Paul Keating,Mwrwan Abubakar,Sirag Elnour

The vSPACE experimental proof-of-concept (PoC) on the TrueElect[Anon][Creds] protocol presents a novel approach to secure, private, and scalable elections, extending the TrueElect and ElectAnon protocols with the integration of AnonCreds SSI (Self-Sovereign Identity). Such a protocol PoC is situated within a Zero-Trust Architecture (ZTA) and leverages confidential computing, continuous authentication, multi-party computation (MPC), and well-architected framework (WAF) principles to address the challenges of cybersecurity, privacy, and trust over IP (ToIP) protection. Employing a Kubernetes confidential cluster within an Enterprise-Scale Landing Zone (ESLZ), vSPACE integrates Distributed Ledger Technology (DLT) for immutable and certifiable audit trails. The Infrastructure as Code (IaC) model ensures rapid deployment, consistent management, and adherence to security standards, making vSPACE a future-proof solution for digital voting systems.

INTERACT · MoDELS · Learning · Prompt · 樣本 ·

2024 年 3 月 8 日

DiffChat: Learning to Chat with Text-to-Image Synthesis Models for Interactive Image Creation

Jiapeng Wang,Chengyu Wang,Tingfeng Cao,Jun Huang,Lianwen Jin

We present DiffChat, a novel method to align Large Language Models (LLMs) to "chat" with prompt-as-input Text-to-Image Synthesis (TIS) models (e.g., Stable Diffusion) for interactive image creation. Given a raw prompt/image and a user-specified instruction, DiffChat can effectively make appropriate modifications and generate the target prompt, which can be leveraged to create the target image of high quality. To achieve this, we first collect an instruction-following prompt engineering dataset named InstructPE for the supervised training of DiffChat. Next, we propose a reinforcement learning framework with the feedback of three core criteria for image creation, i.e., aesthetics, user preference, and content integrity. It involves an action-space dynamic modification technique to obtain more relevant positive samples and harder negative samples during the off-policy sampling. Content integrity is also introduced into the value estimation function for further improvement of produced images. Our method can exhibit superior performance than baseline models and strong competitors based on both automatic and human evaluations, which fully demonstrates its effectiveness.

MoDELS · OD · 估計/估計量 · Markov · 推斷 ·

2024 年 3 月 7 日

Bayesian Inference of Time-Varying Origin-Destination Matrices from Boarding/Alighting Counts for Transit Services

Xiaoxu Chen,Zhanhong Cheng,Lijun Sun

Origin-destination (OD) demand matrices are crucial for transit agencies to design and operate transit systems. This paper presents a novel temporal Bayesian model designed to estimate transit OD matrices at the individual bus-journey level from boarding/alighting counts at bus stops. Our approach begins by modeling the number of alighting passengers at subsequent bus stops, given a boarding stop, through a multinomial distribution parameterized by alighting probabilities. Given the large scale of the problem, we generate alighting probabilities with a latent variable matrix and factorize it into a mapping matrix and a temporal matrix, thereby substantially reducing the number of parameters. To further encode a temporally-smooth structure in the parameters, we impose a Gaussian process prior on the columns of the temporal factor matrix. For model inference, we develop a two-stage algorithm with the Markov chain Monte Carlo (MCMC) method. In the first stage, latent OD matrices are sampled conditional on model parameters using a Metropolis-Hastings sampling algorithm with a Markov model-based proposal distribution. In the second stage, we sample model parameters conditional on latent OD matrices using slice and elliptical slice sampling algorithms. We assess the proposed model using real-world data collected from three bus routes with varying numbers of stops, and the results demonstrate that our model achieves accurate posterior mean estimation and outperforms the widely used iterative proportional fitting (IPF) method. Additionally, our model can provide uncertainty quantification for the OD demand matrices, thus benefiting many downstream planning/operational tasks that require robust decisions.

Analysis · 情感分析 · 數據集 · 可理解性 · 統計量 ·

2024 年 3 月 7 日

MaCmS: Magahi Code-mixed Dataset for Sentiment Analysis

Priya Rani,Gaurav Negi,Theodorus Fransen,John P. McCrae

from arxiv, Lrec-Colin 2024

The present paper introduces new sentiment data, MaCMS, for Magahi-Hindi-English (MHE) code-mixed language, where Magahi is a less-resourced minority language. This dataset is the first Magahi-Hindi-English code-mixed dataset for sentiment analysis tasks. Further, we also provide a linguistics analysis of the dataset to understand the structure of code-mixing and a statistical study to understand the language preferences of speakers with different polarities. With these analyses, we also train baseline models to evaluate the dataset's quality.

MoDELS · Networking · Neural Networks · 可約的 · 相互獨立的 ·

2024 年 3 月 7 日

A Model Hierarchy for Predicting the Flow in Stirred Tanks with Physics-Informed Neural Networks

Veronika Trávníková,Daniel Wolff,Nico Dirkes,Stefanie Elgeti,Eric von Lieres,Marek Behr

from arxiv, 39 pages, 19 figures

This paper explores the potential of Physics-Informed Neural Networks (PINNs) to serve as Reduced Order Models (ROMs) for simulating the flow field within stirred tank reactors (STRs). We solve the two-dimensional stationary Navier-Stokes equations within a geometrically intricate domain and explore methodologies that allow us to integrate additional physical insights into the model. These approaches include imposing the Dirichlet boundary conditions (BCs) strongly and employing domain decomposition (DD), with both overlapping and non-overlapping subdomains. We adapt the Extended Physics-Informed Neural Network (XPINN) approach to solve different sets of equations in distinct subdomains based on the diverse flow characteristics present in each region. Our exploration results in a hierarchy of models spanning various levels of complexity, where the best models exhibit l1 prediction errors of less than 1% for both pressure and velocity. To illustrate the reproducibility of our approach, we track the errors over repeated independent training runs of the best identified model and show its reliability. Subsequently, by incorporating the stirring rate as a parametric input, we develop a fast-to-evaluate model of the flow capable of interpolating across a wide range of Reynolds numbers. Although we exclusively restrict ourselves to STRs in this work, we conclude that the steps taken to obtain the presented model hierarchy can be transferred to other applications.

HTTPS · 代碼 · MoDELS · 語言模型化 · 穩健性 ·

2024 年 3 月 7 日

Evaluation of LLMs on Syntax-Aware Code Fill-in-the-Middle Tasks

Linyuan Gong,Sida Wang,Mostafa Elhoushi,Alvin Cheung

We introduce Syntax-Aware Fill-In-the-Middle (SAFIM), a new benchmark for evaluating Large Language Models (LLMs) on the code Fill-in-the-Middle (FIM) task. This benchmark focuses on syntax-aware completions of program structures such as code blocks and conditional expressions, and includes 17,720 examples from multiple programming languages, sourced from recent code submissions after April 2022 to minimize data contamination. SAFIM provides a robust framework with various prompt designs and novel syntax-aware post-processing techniques, facilitating accurate and fair comparisons across LLMs. Our comprehensive evaluation of 15 LLMs shows that FIM pretraining not only enhances FIM proficiency but also improves Left-to-Right (L2R) inference using LLMs. Our findings challenge conventional beliefs and suggest that pretraining methods and data quality have more impact than model size. SAFIM thus serves as a foundational platform for future research in effective pretraining strategies for code LLMs. The evaluation toolkit and dataset are available at //github.com/gonglinyuan/safim, and the leaderboard is available at //safimbenchmark.com.

LIDAR · Processing（編程語言） · state-of-the-art · 3D · CARS ·

2024 年 3 月 6 日

Multi-Object Tracking with Camera-LiDAR Fusion for Autonomous Driving

Riccardo Pieroni,Simone Specchia,Matteo Corno,Sergio Matteo Savaresi

from arxiv, Published at IEEE European Control Conference 2024

This paper presents a novel multi-modal Multi-Object Tracking (MOT) algorithm for self-driving cars that combines camera and LiDAR data. Camera frames are processed with a state-of-the-art 3D object detector, whereas classical clustering techniques are used to process LiDAR observations. The proposed MOT algorithm comprises a three-step association process, an Extended Kalman filter for estimating the motion of each detected dynamic obstacle, and a track management phase. The EKF motion model requires the current measured relative position and orientation of the observed object and the longitudinal and angular velocities of the ego vehicle as inputs. Unlike most state-of-the-art multi-modal MOT approaches, the proposed algorithm does not rely on maps or knowledge of the ego global pose. Moreover, it uses a 3D detector exclusively for cameras and is agnostic to the type of LiDAR sensor used. The algorithm is validated both in simulation and with real-world data, with satisfactory results.