成人艳情一二三区按摩_91成人精品爽啪在线观看_看全色黄大色黄大片女爽一黄_韩国一区二区熟睡人妻视频_欧美激情在线播放一区二区_亚洲日韩国产另类_4010久精品视在线观看视频

In the realm of Multimodal Large Language Models (MLLMs), vision-language connector plays a crucial role to link the pre-trained vision encoders with Large Language Models (LLMs). Despite its importance, the vision-language connector has been relatively less explored. In this study, we aim to propose a strong vision-language connector that enables MLLMs to achieve high accuracy while maintain low computation cost. We first reveal the existence of the visual anchors in Vision Transformer and propose a cost-effective search algorithm to extract them. Building on these findings, we introduce the Anchor Former (AcFormer), a novel vision-language connector designed to leverage the rich prior knowledge obtained from these visual anchors during pretraining, guiding the aggregation of information. Through extensive experimentation, we demonstrate that the proposed method significantly reduces computational costs by nearly two-thirds compared with baseline, while simultaneously outperforming baseline methods. This highlights the effectiveness and efficiency of AcFormer. Codes are available at //github.com/liuhaogeng/Anchor-Former.

相關內容

anchor

關注 0

Performer · Networking · 圖 · 結點 · Agent ·

2024 年 12 月 15 日

Enhancing Multiagent Genetic Network Programming Performance Using Search Space Reduction

Ali Kohan,Mohamad Roshanzamir,Roohallah Alizadehsani

Genetic Network Programming (GNP) is an evolutionary algorithm that extends Genetic Programming (GP). It is typically used in agent control problems. In contrast to GP, which employs a tree structure, GNP utilizes a directed graph structure. During the evolutionary process, the connections between nodes change to discover the optimal strategy. Due to the large number of node connections, GNP has a large search space, making it challenging to identify an appropriate graph structure. One way to reduce this search space is by utilizing simplified operators that restrict the changeable node connections to those participating in the fitness function. However, this method has not been applied to GNP structures that use separate graphs for each agent, such as situation-based GNP (SBGNP). This paper proposes a method to apply simplified operators to SBGNP. To evaluate the performance of this method, we tested it on the Tileworld benchmark, where the algorithm demonstrated improvements in average fitness.

圖 · 穩健性 · Neural Networks · Networking · 圖形處理器 ·

2024 年 12 月 14 日

Improving Graph Neural Networks via Adversarial Robustness Evaluation

Yongyu Wang

Graph Neural Networks (GNNs) are currently one of the most powerful types of neural network architectures. Their advantage lies in the ability to leverage both the graph topology, which represents the relationships between samples, and the features of the samples themselves. However, the given graph topology often contains noisy edges, and GNNs are vulnerable to noise in the graph structure. This issue remains unresolved. In this paper, we propose using adversarial robustness evaluation to select a small subset of robust nodes that are less affected by noise. We then only feed the features of these robust nodes, along with the KNN graph constructed from these nodes, into the GNN for classification. Additionally, we compute the centroids for each class. For the remaining non-robust nodes, we assign them to the class whose centroid is closest to them. Experimental results show that this method significantly improves the accuracy of GNNs.

通道 · 優化器 · Neural Networks · 情景 · Networking ·

2024 年 12 月 13 日

Single Channel EEG Based Insomnia Identification Without Sleep Stage Annotations

Chan-Yun Yang,Nilantha Premakumara,Hooman Samani,Chinthaka Premachandra

from arxiv, 30 Pages, 9 figures and 12 tables

This paper proposes a new approach to identifying patients with insomnia using a single EEG channel, without the need for sleep stage annotation. Data preprocessing, feature extraction, feature selection, and classification techniques are used to automatically detect insomnia based on features extracted from spectral and temporal domains, including relative power in the delta, sigma, beta and gamma bands, total power, absolute slow wave power, power ratios, mean, zero crossing rate, mobility, and complexity. A Pearson correlation coefficient, t-test, p-value, and two rules are used to select the optimal set of features for accurately classifying insomnia patients and rejecting negatively affecting features. Classification schemes including a general artificial neural network, convolutional neural network, and support vector machine are applied to the optimal feature set to distinguish between insomnia patients and healthy subjects. The performance of the model is validated using 50 insomnia patients and 50 healthy subjects, with the Fp2 channel and 1D-CNN classifier achieving the highest accuracy and Cohen's kappa coefficient at 97.85% and 94.15%, respectively. The developed model has the potential to simplify current sleep monitoring systems and enable in-home ambulatory monitoring.

縮放 · MoDELS · Better · 推斷 · 潛在 ·

2024 年 12 月 13 日

Byte Latent Transformer: Patches Scale Better Than Tokens

Artidoro Pagnoni,Ram Pasunuru,Pedro Rodriguez,John Nguyen,Benjamin Muller,Margaret Li,Chunting Zhou,Lili Yu,Jason Weston,Luke Zettlemoyer,Gargi Ghosh,Mike Lewis,Ari Holtzman,Srinivasan Iyer

We introduce the Byte Latent Transformer (BLT), a new byte-level LLM architecture that, for the first time, matches tokenization-based LLM performance at scale with significant improvements in inference efficiency and robustness. BLT encodes bytes into dynamically sized patches, which serve as the primary units of computation. Patches are segmented based on the entropy of the next byte, allocating more compute and model capacity where increased data complexity demands it. We present the first FLOP controlled scaling study of byte-level models up to 8B parameters and 4T training bytes. Our results demonstrate the feasibility of scaling models trained on raw bytes without a fixed vocabulary. Both training and inference efficiency improve due to dynamically selecting long patches when data is predictable, along with qualitative improvements on reasoning and long tail generalization. Overall, for fixed inference costs, BLT shows significantly better scaling than tokenization-based models, by simultaneously growing both patch and model size.

蒙特卡羅 · Lipschitz · Continuity · 控制器 · 矩 ·

2024 年 12 月 12 日

Langevin Monte Carlo Beyond Lipschitz Gradient Continuity

Matej Benko,Iwona Chlebicka,J?rgen Endal,B?a?ej Miasojedow

from arxiv, To appear in Proceedings of the AAAI Conference on Artificial Intelligence (AAAI-25)

We present a significant advancement in the field of Langevin Monte Carlo (LMC) methods by introducing the Inexact Proximal Langevin Algorithm (IPLA). This novel algorithm broadens the scope of problems that LMC can effectively address while maintaining controlled computational costs. IPLA extends LMC's applicability to potentials that are convex, strongly convex in the tails, and exhibit polynomial growth, beyond the conventional $L$-smoothness assumption. Moreover, we extend LMC's applicability to super-quadratic potentials and offer improved convergence rates over existing algorithms. Additionally, we provide bounds on all moments of the Markov chain generated by IPLA, enhancing its analytical robustness.

Vision · 變換 · Networking · state-of-the-art · Performer ·

2024 年 12 月 12 日

Vision Transformers for Efficient Indoor Pathloss Radio Map Prediction

Edvard Ghukasyan,Hrant Khachatrian,Rafayel Mkrtchyan,Theofanis P. Raptis

from arxiv, Work partly supported by the RA Science Committee grant No. 22rl-052 (DISTAL) and the EU under Italian National Recovery and Resilience Plan of NextGenerationEU on "Telecommunications of the Future" (PE00000001 - program "RESTART")

Vision Transformers (ViTs) have demonstrated remarkable success in achieving state-of-the-art performance across various image-based tasks and beyond. In this study, we employ a ViT-based neural network to address the problem of indoor pathloss radio map prediction. The network's generalization ability is evaluated across diverse settings, including unseen buildings, frequencies, and antennas with varying radiation patterns. By leveraging extensive data augmentation techniques and pretrained DINOv2 weights, we achieve promising results, even under the most challenging scenarios.

Performer · GPS · MoDELS · 優化器 · 層 ·

2024 年 12 月 12 日

Bayesian Optimization via Continual Variational Last Layer Training

Paul Brunzema,Mikkel Jordahn,John Willes,Sebastian Trimpe,Jasper Snoek,James Harrison

Gaussian Processes (GPs) are widely seen as the state-of-the-art surrogate models for Bayesian optimization (BO) due to their ability to model uncertainty and their performance on tasks where correlations are easily captured (such as those defined by Euclidean metrics) and their ability to be efficiently updated online. However, the performance of GPs depends on the choice of kernel, and kernel selection for complex correlation structures is often difficult or must be made bespoke. While Bayesian neural networks (BNNs) are a promising direction for higher capacity surrogate models, they have so far seen limited use due to poor performance on some problem types. In this paper, we propose an approach which shows competitive performance on many problem types, including some that BNNs typically struggle with. We build on variational Bayesian last layers (VBLLs), and connect training of these models to exact conditioning in GPs. We exploit this connection to develop an efficient online training algorithm that interleaves conditioning and optimization. Our findings suggest that VBLL networks significantly outperform GPs and other BNN architectures on tasks with complex input correlations, and match the performance of well-tuned GPs on established benchmark tasks.

MoDELS · Performer · HTTPS · 監督 · 組合性 ·

2023 年 12 月 4 日

Data Management For Large Language Models: A Survey

Zige Wang,Wanjun Zhong,Yufei Wang,Qi Zhu,Fei Mi,Baojun Wang,Lifeng Shang,Xin Jiang,Qun Liu

from arxiv, Work in progress

Data plays a fundamental role in the training of Large Language Models (LLMs). Effective data management, particularly in the formulation of a well-suited training dataset, holds significance for enhancing model performance and improving training efficiency during pretraining and supervised fine-tuning phases. Despite the considerable importance of data management, the current research community still falls short in providing a systematic analysis of the rationale behind management strategy selection, its consequential effects, methodologies for evaluating curated datasets, and the ongoing pursuit of improved strategies. Consequently, the exploration of data management has attracted more and more attention among the research community. This survey provides a comprehensive overview of current research in data management within both the pretraining and supervised fine-tuning stages of LLMs, covering various noteworthy aspects of data management strategy design: data quantity, data quality, domain/task composition, etc. Looking toward the future, we extrapolate existing challenges and outline promising directions for development in this field. Therefore, this survey serves as a guiding resource for practitioners aspiring to construct powerful LLMs through effective data management practices. The collection of the latest papers is available at //github.com/ZigeW/data_management_LLM.

圖形處理器 · 圖 · Neural Networks · Networking · Performer ·

2021 年 2 月 13 日

How Framelets Enhance Graph Neural Networks

Xuebin Zheng,Bingxin Zhou,Junbin Gao,Yu Guang Wang,Pietro Lio,Ming Li,Guido Montufar

from arxiv, 24 pages, 17 figures, 6 tables

This paper presents a new approach for assembling graph neural networks based on framelet transforms. The latter provides a multi-scale representation for graph-structured data. With the framelet system, we can decompose the graph feature into low-pass and high-pass frequencies as extracted features for network training, which then defines a framelet-based graph convolution. The framelet decomposition naturally induces a graph pooling strategy by aggregating the graph feature into low-pass and high-pass spectra, which considers both the feature values and geometry of the graph data and conserves the total information. The graph neural networks with the proposed framelet convolution and pooling achieve state-of-the-art performance in many types of node and graph prediction tasks. Moreover, we propose shrinkage as a new activation for the framelet convolution, which thresholds the high-frequency information at different scales. Compared to ReLU, shrinkage in framelet convolution improves the graph neural network model in terms of denoising and signal compression: noises in both node and structure can be significantly reduced by accurately cutting off the high-pass coefficients from framelet decomposition, and the signal can be compressed to less than half its original size with the prediction performance well preserved.

模式崩潰 · 對抗自編碼 · 自編碼器 · 峰值 · Better ·

2018 年 3 月 23 日

Generative Adversarial Autoencoder Networks

Ngoc-Trung Tran,Tuan-Anh Bui,Ngai-Man Cheung

We introduce an effective model to overcome the problem of mode collapse when training Generative Adversarial Networks (GAN). Firstly, we propose a new generator objective that finds it better to tackle mode collapse. And, we apply an independent Autoencoders (AE) to constrain the generator and consider its reconstructed samples as "real" samples to slow down the convergence of discriminator that enables to reduce the gradient vanishing problem and stabilize the model. Secondly, from mappings between latent and data spaces provided by AE, we further regularize AE by the relative distance between the latent and data samples to explicitly prevent the generator falling into mode collapse setting. This idea comes when we find a new way to visualize the mode collapse on MNIST dataset. To the best of our knowledge, our method is the first to propose and apply successfully the relative distance of latent and data samples for stabilizing GAN. Thirdly, our proposed model, namely Generative Adversarial Autoencoder Networks (GAAN), is stable and has suffered from neither gradient vanishing nor mode collapse issues, as empirically demonstrated on synthetic, MNIST, MNIST-1K, CelebA and CIFAR-10 datasets. Experimental results show that our method can approximate well multi-modal distribution and achieve better results than state-of-the-art methods on these benchmark datasets. Our model implementation is published here: //github.com/tntrung/gaan