Trending Papers

8

GitHub 4.46k arXiv Page

Submitted by

amael-apple

Sharp Monocular View Synthesis in Less Than a Second

SHARP synthesizes photorealistic views from a single image using a 3D Gaussian representation, achieving state-of-the-art results with rapid processing.

Apple · Dec 11, 2025

8

GitHub 4.46k arXiv Page

Submitted by

hao-li

Agent READMEs: An Empirical Study of Context Files for Agentic Coding

Agentic coding tools receive goals written in natural language as input, break them down into specific tasks, and write or execute the actual code with minimal human intervention. Central to this process are agent context files ("READMEs for agents") that provide persistent, project-level instructions. In this paper, we conduct the first large-scale empirical study of 2,303 agent context files from 1,925 repositories to characterize their structure, maintenance, and content. We find that these files are not static documentation but complex, difficult-to-read artifacts that evolve like configuration code, maintained through frequent, small additions. Our content analysis of 16 instruction types shows that developers prioritize functional context, such as build and run commands (62.3%), implementation details (69.9%), and architecture (67.7%). We also identify a significant gap: non-functional requirements like security (14.5%) and performance (14.5%) are rarely specified. These findings indicate that while developers use context files to make agents functional, they provide few guardrails to ensure that agent-written code is secure or performant, highlighting the need for improved tooling and practices.

11 authors

· Published on Nov 17, 2025

Submitted by

hao-li

Agent READMEs: An Empirical Study of Context Files for Agentic Coding

Agentic coding tools receive goals written in natural language as input, break them down into specific tasks, and write or execute the actual code with minimal human intervention. Central to this process are agent context files ("READMEs for agents") that provide persistent, project-level instructions. In this paper, we conduct the first large-scale empirical study of 2,303 agent context files from 1,925 repositories to characterize their structure, maintenance, and content. We find that these files are not static documentation but complex, difficult-to-read artifacts that evolve like configuration code, maintained through frequent, small additions. Our content analysis of 16 instruction types shows that developers prioritize functional context, such as build and run commands (62.3%), implementation details (69.9%), and architecture (67.7%). We also identify a significant gap: non-functional requirements like security (14.5%) and performance (14.5%) are rarely specified. These findings indicate that while developers use context files to make agents functional, they provide few guardrails to ensure that agent-written code is secure or performant, highlighting the need for improved tooling and practices.

11 authors

· Nov 17, 2025

Submitted by

unilm

VibeVoice Technical Report

VibeVoice synthesizes long-form multi-speaker speech using next-token diffusion and a highly efficient continuous speech tokenizer, achieving superior performance and fidelity.

Microsoft Research · Published on Aug 26, 2025

135

GitHub 18.8k arXiv Page

Submitted by

unilm

VibeVoice Technical Report

VibeVoice synthesizes long-form multi-speaker speech using next-token diffusion and a highly efficient continuous speech tokenizer, achieving superior performance and fidelity.

Microsoft Research · Aug 26, 2025

135

GitHub 18.8k arXiv Page

Submitted by

andito

SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

SmolDocling is a compact vision-language model that performs end-to-end document conversion with robust performance across various document types using 256M parameters and a new markup format.

13 authors

· Published on Mar 14, 2025

120

GitHub 47.4k arXiv Page

Submitted by

andito

SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

SmolDocling is a compact vision-language model that performs end-to-end document conversion with robust performance across various document types using 256M parameters and a new markup format.

13 authors

· Mar 14, 2025

120

GitHub 47.4k arXiv Page

Submitted by

taesiri

WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling

WorldPlay is a streaming video diffusion model that achieves real-time, interactive world modeling with long-term geometric consistency by using a Dual Action Representation, Reconstituted Context Memory, and Context Forcing.

10 authors

· Published on Dec 16, 2025

61

GitHub 623 arXiv Page

Submitted by

taesiri

WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling

WorldPlay is a streaming video diffusion model that achieves real-time, interactive world modeling with long-term geometric consistency by using a Dual Action Representation, Reconstituted Context Memory, and Context Forcing.

10 authors

· Dec 16, 2025

61

GitHub 623 arXiv Page

Submitted by

Cxxs

Decoupled DMD: CFG Augmentation as the Spear, Distribution Matching as the Shield

The study reveals that in text-to-image generation, CFG Augmentation is the primary driver of few-step distillation in Distribution Matching Distillation (DMD), while the distribution matching term acts as a regularizer.

Tongyi-MAI · Published on Nov 27, 2025

GitHub 7.57k arXiv Page

Submitted by

Cxxs

Decoupled DMD: CFG Augmentation as the Spear, Distribution Matching as the Shield

The study reveals that in text-to-image generation, CFG Augmentation is the primary driver of few-step distillation in Distribution Matching Distillation (DMD), while the distribution matching term acts as a regularizer.

Tongyi-MAI · Nov 27, 2025

GitHub 7.57k arXiv Page

Submitted by

Paper99

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Z-Image, a 6B-parameter Scalable Single-Stream Diffusion Transformer (S3-DiT) model, achieves high-performance image generation with reduced computational cost, offering sub-second inference and compatibility with consumer hardware.

Tongyi-MAI · Published on Nov 27, 2025

205

GitHub 7.56k arXiv Page

Submitted by

Paper99

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Z-Image, a 6B-parameter Scalable Single-Stream Diffusion Transformer (S3-DiT) model, achieves high-performance image generation with reduced computational cost, offering sub-second inference and compatibility with consumer hardware.

Tongyi-MAI · Nov 27, 2025

205

GitHub 7.56k arXiv Page

Self-Supervised Prompt Optimization

A self-supervised framework optimizes prompts for both closed and open-ended tasks by evaluating LLM outputs without external references, reducing costs and required data.

9 authors

· Published on Feb 7, 2025

GitHub 61.1k arXiv Page

Self-Supervised Prompt Optimization

A self-supervised framework optimizes prompts for both closed and open-ended tasks by evaluating LLM outputs without external references, reducing costs and required data.

9 authors

· Feb 7, 2025

GitHub 61.1k arXiv Page

Submitted by

taesiri

DeepCode: Open Agentic Coding

DeepCode, a fully autonomous framework, addresses the challenges of document-to-codebase synthesis by optimizing information flow through source compression, structured indexing, knowledge injection, and error correction, achieving state-of-the-art performance and surpassing human experts.

5 authors

· Published on Dec 8, 2025

30

Submitted by

taesiri

DeepCode: Open Agentic Coding

DeepCode, a fully autonomous framework, addresses the challenges of document-to-codebase synthesis by optimizing information flow through source compression, structured indexing, knowledge injection, and error correction, achieving state-of-the-art performance and surpassing human experts.

5 authors

· Dec 8, 2025

30

Submitted by

akhaliq

Efficient Memory Management for Large Language Model Serving with PagedAttention

PagedAttention algorithm and vLLM system enhance the throughput of large language models by efficiently managing memory and reducing waste in the key-value cache.

9 authors

· Published on Sep 12, 2023

GitHub 65.9k arXiv Page

Submitted by

akhaliq

Efficient Memory Management for Large Language Model Serving with PagedAttention

PagedAttention algorithm and vLLM system enhance the throughput of large language models by efficiently managing memory and reducing waste in the key-value cache.

9 authors

· Sep 12, 2023

GitHub 65.9k arXiv Page

Submitted by

taesiri

Step-GUI Technical Report

A self-evolving training pipeline with the Calibrated Step Reward System and GUI-MCP protocol improve GUI automation efficiency, accuracy, and privacy in real-world scenarios.

StepFun · Published on Dec 17, 2025

119

GitHub 1.65k arXiv Page

Submitted by

taesiri

Step-GUI Technical Report

A self-evolving training pipeline with the Calibrated Step Reward System and GUI-MCP protocol improve GUI automation efficiency, accuracy, and privacy in real-world scenarios.

StepFun · Dec 17, 2025

119

GitHub 1.65k arXiv Page

Submitted by

taesiri

SAM 3: Segment Anything with Concepts

Segment Anything Model 3 achieves state-of-the-art performance in promptable concept segmentation and tracking by leveraging a unified model architecture with decoupled recognition and localization.

AI at Meta · Published on Nov 20, 2025

118

GitHub 6.28k arXiv Page

Submitted by

taesiri

SAM 3: Segment Anything with Concepts

Segment Anything Model 3 achieves state-of-the-art performance in promptable concept segmentation and tracking by leveraging a unified model architecture with decoupled recognition and localization.

AI at Meta · Nov 20, 2025

118

GitHub 6.28k arXiv Page

Submitted by

akhaliq

FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

FunAudioLLM enhances voice interactions by integrating SenseVoice for multilingual speech recognition, emotion detection, and audio event detection with CosyVoice for natural speech generation across languages, timbres, and styles.

1 authors

· Published on Jul 4, 2024

40

GitHub 18.2k arXiv Page

Submitted by

akhaliq

FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

FunAudioLLM enhances voice interactions by integrating SenseVoice for multilingual speech recognition, emotion detection, and audio event detection with CosyVoice for natural speech generation across languages, timbres, and styles.

1 authors

· Jul 4, 2024

40

GitHub 18.2k arXiv Page

AI-Trader: Benchmarking Autonomous Agents in Real-Time Financial Markets

AI-Trader evaluates the performance of large language models in real-world financial markets, highlighting their limitations in trading and risk management.

6 authors

· Published on Dec 1, 2025

GitHub 10.2k arXiv Page

AI-Trader: Benchmarking Autonomous Agents in Real-Time Financial Markets

AI-Trader evaluates the performance of large language models in real-world financial markets, highlighting their limitations in trading and risk management.

6 authors

· Dec 1, 2025

GitHub 10.2k arXiv Page

Submitted by

taesiri

PersonaLive! Expressive Portrait Image Animation for Live Streaming

PersonaLive is a diffusion-based framework for real-time portrait animation that enhances speed and efficiency through multi-stage training, hybrid implicit signals, appearance distillation, and autoregressive micro-chunk streaming.

GVC Lab at Great Bay University · Published on Dec 12, 2025

GitHub 631 arXiv Page

Submitted by

taesiri

PersonaLive! Expressive Portrait Image Animation for Live Streaming

PersonaLive is a diffusion-based framework for real-time portrait animation that enhances speed and efficiency through multi-stage training, hybrid implicit signals, appearance distillation, and autoregressive micro-chunk streaming.

GVC Lab at Great Bay University · Dec 12, 2025

GitHub 631 arXiv Page

Submitted by

taesiri

PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Ultra-Compact Vision-Language Model

PaddleOCR-VL, a vision-language model combining NaViT-style dynamic resolution and ERNIE, achieves state-of-the-art performance in document parsing and element recognition with high efficiency.

PaddlePaddle · Published on Oct 16, 2025

107

GitHub 66.6k arXiv Page

Submitted by

taesiri

PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Ultra-Compact Vision-Language Model

PaddleOCR-VL, a vision-language model combining NaViT-style dynamic resolution and ERNIE, achieves state-of-the-art performance in document parsing and element recognition with high efficiency.

PaddlePaddle · Oct 16, 2025

107

GitHub 66.6k arXiv Page

Submitted by

LiheYoung

In Pursuit of Pixel Supervision for Visual Pre-training

Pixio, an enhanced masked autoencoder, demonstrates competitive performance across various downstream tasks using pixel-space self-supervised learning, outperforming latent-space approaches.

8 authors

· Published on Dec 17, 2025

GitHub 155 arXiv Page

Submitted by

LiheYoung

In Pursuit of Pixel Supervision for Visual Pre-training

Pixio, an enhanced masked autoencoder, demonstrates competitive performance across various downstream tasks using pixel-space self-supervised learning, outperforming latent-space approaches.

8 authors

· Dec 17, 2025

GitHub 155 arXiv Page

Submitted by

taesiri

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

MinerU2.5, a 1.2B-parameter document parsing vision-language model, achieves state-of-the-art recognition accuracy with computational efficiency through a coarse-to-fine parsing strategy.

61 authors

· Published on Sep 26, 2025

139

Submitted by

taesiri

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

MinerU2.5, a 1.2B-parameter document parsing vision-language model, achieves state-of-the-art recognition accuracy with computational efficiency through a coarse-to-fine parsing strategy.

61 authors

· Sep 26, 2025

139

Submitted by

wanderkid

MinerU: An Open-Source Solution for Precise Document Content Extraction

MinerU is an open-source tool that enhances document content extraction using fine-tuned models and pre/postprocessing rules across diverse document types.

18 authors

· Published on Sep 27, 2024

Submitted by

wanderkid

MinerU: An Open-Source Solution for Precise Document Content Extraction

MinerU is an open-source tool that enhances document content extraction using fine-tuned models and pre/postprocessing rules across diverse document types.

18 authors

· Sep 27, 2024

Submitted by

taesiri

Memory in the Age of AI Agents

This survey provides an updated overview of agent memory research, distinguishing its forms, functions, and dynamics, and highlights emerging research directions.

47 authors

· Published on Dec 15, 2025

103

GitHub 399 arXiv Page

Submitted by

taesiri

Memory in the Age of AI Agents

This survey provides an updated overview of agent memory research, distinguishing its forms, functions, and dynamics, and highlights emerging research directions.

47 authors

· Dec 15, 2025

103

GitHub 399 arXiv Page

Submitted by

FrancisRing

FlashPortrait: 6x Faster Infinite Portrait Animation with Adaptive Latent Prediction

FlashPortrait is a diffusion-based video transformer for long-portrait animation that ensures ID consistency and achieves 6x acceleration through a dynamic sliding-window scheme and higher-order latent derivatives.

Fudan University · Published on Dec 18, 2025

GitHub 98 arXiv Page

Submitted by

FrancisRing

FlashPortrait: 6x Faster Infinite Portrait Animation with Adaptive Latent Prediction

FlashPortrait is a diffusion-based video transformer for long-portrait animation that ensures ID consistency and achieves 6x acceleration through a dynamic sliding-window scheme and higher-order latent derivatives.

Fudan University · Dec 18, 2025

GitHub 98 arXiv Page

LightRAG: Simple and Fast Retrieval-Augmented Generation

LightRAG improves Retrieval-Augmented Generation by integrating graph structures for enhanced contextual awareness and efficient information retrieval, achieving better accuracy and response times.

5 authors

· Published on Oct 8, 2024

GitHub 26.2k arXiv Page

LightRAG: Simple and Fast Retrieval-Augmented Generation

LightRAG improves Retrieval-Augmented Generation by integrating graph structures for enhanced contextual awareness and efficient information retrieval, achieving better accuracy and response times.

5 authors

· Oct 8, 2024

GitHub 26.2k arXiv Page

Submitted by

rmurthy

Promptomatix: An Automatic Prompt Optimization Framework for Large Language Models

Promptomatix automates prompt optimization for Large Language Models, improving performance and efficiency across various tasks.

9 authors

· Published on Jul 17, 2025

17

GitHub 799 arXiv Page

Submitted by

rmurthy

Promptomatix: An Automatic Prompt Optimization Framework for Large Language Models

Promptomatix automates prompt optimization for Large Language Models, improving performance and efficiency across various tasks.

9 authors

· Jul 17, 2025

17

GitHub 799 arXiv Page

Submitted by

Wayne-King

Next-Embedding Prediction Makes Strong Vision Learners

Generative pretraining using next embedding prediction outperforms traditional self-supervised methods in visual learning tasks, achieving high accuracy on ImageNet and effective transfer to semantic segmentation.

SixAILab · Published on Dec 18, 2025

62

GitHub 87 arXiv Page

Submitted by

Wayne-King

Next-Embedding Prediction Makes Strong Vision Learners

Generative pretraining using next embedding prediction outperforms traditional self-supervised methods in visual learning tasks, achieving high accuracy on ImageNet and effective transfer to semantic segmentation.

SixAILab · Dec 18, 2025

62

GitHub 87 arXiv Page

Submitted by

AdinaY

SCAIL: Towards Studio-Grade Character Animation via In-Context Learning of 3D-Consistent Pose Representations

SCAIL framework improves character animation by using a novel 3D pose representation and a diffusion-transformer architecture with full-context pose injection, achieving studio-grade quality and realism.

Z.ai · Published on Dec 5, 2025

19

GitHub 436 arXiv Page

Submitted by

AdinaY

SCAIL: Towards Studio-Grade Character Animation via In-Context Learning of 3D-Consistent Pose Representations

SCAIL framework improves character animation by using a novel 3D pose representation and a diffusion-transformer architecture with full-context pose injection, achieving studio-grade quality and realism.

Z.ai · Dec 5, 2025

19

GitHub 436 arXiv Page

Submitted by

akhaliq

LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

LlamaFactory is a unified framework enabling efficient fine-tuning of large language models across various tasks using a web-based user interface.

5 authors

· Published on Mar 20, 2024

175

GitHub 64.3k arXiv Page

Submitted by

akhaliq

LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

LlamaFactory is a unified framework enabling efficient fine-tuning of large language models across various tasks using a web-based user interface.

5 authors

· Mar 20, 2024

175

GitHub 64.3k arXiv Page

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

IndexTTS, an enhanced text-to-speech system combining XTTS and Tortoise models, offers improved naturalness, enhanced voice cloning, and controllable usage through hybrid character-pinyin modeling and optimized vector quantization.

5 authors

· Published on Feb 8, 2025

GitHub 16.9k arXiv Page

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

IndexTTS, an enhanced text-to-speech system combining XTTS and Tortoise models, offers improved naturalness, enhanced voice cloning, and controllable usage through hybrid character-pinyin modeling and optimized vector quantization.

5 authors

· Feb 8, 2025

GitHub 16.9k arXiv Page

Submitted by

Yhmeng1106

The World is Your Canvas: Painting Promptable Events with Reference Images, Trajectories, and Text

WorldCanvas generates coherent, controllable world events using a multimodal framework that integrates text, trajectories, and reference images.

Ant Group · Published on Dec 18, 2025

23

GitHub 74 arXiv Page

Submitted by

Yhmeng1106

The World is Your Canvas: Painting Promptable Events with Reference Images, Trajectories, and Text

WorldCanvas generates coherent, controllable world events using a multimodal framework that integrates text, trajectories, and reference images.

Ant Group · Dec 18, 2025

23

GitHub 74 arXiv Page

Submitted by

taesiri

LongCat-Video Technical Report

LongCat-Video, a 13.6B parameter video generation model based on the Diffusion Transformer framework, excels in efficient and high-quality long video generation across multiple tasks using unified architecture, coarse-to-fine generation, and block sparse attention.

LongCat · Published on Oct 25, 2025

GitHub 1.58k arXiv Page

Submitted by

taesiri

LongCat-Video Technical Report

LongCat-Video, a 13.6B parameter video generation model based on the Diffusion Transformer framework, excels in efficient and high-quality long video generation across multiple tasks using unified architecture, coarse-to-fine generation, and block sparse attention.

LongCat · Oct 25, 2025

GitHub 1.58k arXiv Page

Submitted by

Snyhlxde

Fast and Accurate Causal Parallel Decoding using Jacobi Forcing

Jacobi Forcing is a progressive distillation method that enables efficient parallel decoding of transformer-based models while maintaining performance, significantly reducing inference latency.

8 authors

· Published on Dec 16, 2025

39

GitHub 141 arXiv Page

Submitted by

Snyhlxde

Fast and Accurate Causal Parallel Decoding using Jacobi Forcing

Jacobi Forcing is a progressive distillation method that enables efficient parallel decoding of transformer-based models while maintaining performance, significantly reducing inference latency.

8 authors

· Dec 16, 2025

39

GitHub 141 arXiv Page

Submitted by

akhaliq

Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory

Mem0, a memory-centric architecture with graph-based memory, enhances long-term conversational coherence in LLMs by efficiently extracting, consolidating, and retrieving information, outperforming existing memory systems in terms of accuracy and computational efficiency.

5 authors

· Published on Apr 28, 2025

GitHub 44.5k arXiv Page

Submitted by

akhaliq

Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory

Mem0, a memory-centric architecture with graph-based memory, enhances long-term conversational coherence in LLMs by efficiently extracting, consolidating, and retrieving information, outperforming existing memory systems in terms of accuracy and computational efficiency.

5 authors

· Apr 28, 2025

GitHub 44.5k arXiv Page

Submitted by

Weiyun1025

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

InternVL3 is a multimodal pre-trained language model that jointly learns from both multimodal data and text, improving performance and scalability through advanced techniques and setting a new state-of-the-art in multimodal tasks.

47 authors

· Published on Apr 14, 2025

306

arXiv Page

Submitted by

Weiyun1025

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

InternVL3 is a multimodal pre-trained language model that jointly learns from both multimodal data and text, improving performance and scalability through advanced techniques and setting a new state-of-the-art in multimodal tasks.

47 authors

· Apr 14, 2025

306

arXiv Page

Submitted by

MapleF9

Towards Scalable Pre-training of Visual Tokenizers for Generation

A unified visual tokenizer pre-training framework (VTP) improves generative performance by optimizing image-text contrastive, self-supervised, and reconstruction losses, leading to better scaling properties and higher zero-shot accuracy and faster convergence.

MiniMax · Published on Dec 15, 2025

89

GitHub 236 arXiv Page

Submitted by

MapleF9

Towards Scalable Pre-training of Visual Tokenizers for Generation

A unified visual tokenizer pre-training framework (VTP) improves generative performance by optimizing image-text contrastive, self-supervised, and reconstruction losses, leading to better scaling properties and higher zero-shot accuracy and faster convergence.

MiniMax · Dec 15, 2025

89

GitHub 236 arXiv Page

OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation

A novel GPT-based model, OmniFlatten, enables real-time natural full-duplex spoken dialogue through a multi-stage post-training technique that integrates speech and text without altering the original model's architecture.

9 authors

· Published on Oct 23, 2024

GitHub 51.3k arXiv Page

OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation

A novel GPT-based model, OmniFlatten, enables real-time natural full-duplex spoken dialogue through a multi-stage post-training technique that integrates speech and text without altering the original model's architecture.

9 authors

· Oct 23, 2024

GitHub 51.3k arXiv Page

From Benchmarks to Business Impact: Deploying IBM Generalist Agent in Enterprise Production

IBM's CUGA, a generalist agent with a hierarchical planner-executor architecture, demonstrates state-of-the-art performance in academic benchmarks and shows potential for enterprise adoption in business-process-outsourcing, addressing scalability, auditability, safety, and governance.

12 authors

· Published on Oct 27, 2025

5

GitHub 493 arXiv Page

From Benchmarks to Business Impact: Deploying IBM Generalist Agent in Enterprise Production

IBM's CUGA, a generalist agent with a hierarchical planner-executor architecture, demonstrates state-of-the-art performance in academic benchmarks and shows potential for enterprise adoption in business-process-outsourcing, addressing scalability, auditability, safety, and governance.

12 authors

· Oct 27, 2025

5

GitHub 493 arXiv Page

TradingAgents: Multi-Agents LLM Financial Trading Framework

A multi-agent framework using large language models for stock trading simulates real-world trading firms, improving performance metrics like cumulative returns and Sharpe ratio.

4 authors

· Published on Dec 28, 2024

14

GitHub 26.8k arXiv Page

TradingAgents: Multi-Agents LLM Financial Trading Framework

A multi-agent framework using large language models for stock trading simulates real-world trading firms, improving performance metrics like cumulative returns and Sharpe ratio.

4 authors

· Dec 28, 2024

14

GitHub 26.8k arXiv Page

Submitted by

xw-eric

Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents

Agent S2, a compositional framework using Mixture-of-Grounding and Proactive Hierarchical Planning, achieves state-of-the-art performance in computer use automation across various benchmarks and operating systems.

Simular · Published on Apr 1, 2025

Submitted by

xw-eric

Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents

Agent S2, a compositional framework using Mixture-of-Grounding and Proactive Hierarchical Planning, achieves state-of-the-art performance in computer use automation across various benchmarks and operating systems.

Simular · Apr 1, 2025

Submitted by

xw-eric

Agent S: An Open Agentic Framework that Uses Computers Like a Human

Agent S, a framework for autonomous GUI interactions, enhances task automation with experience-augmented hierarchical planning and Multimodal Large Language Models.

6 authors

· Published on Oct 10, 2024

GitHub 8.95k arXiv Page

Submitted by

xw-eric

Agent S: An Open Agentic Framework that Uses Computers Like a Human

Agent S, a framework for autonomous GUI interactions, enhances task automation with experience-augmented hierarchical planning and Multimodal Large Language Models.

6 authors

· Oct 10, 2024

GitHub 8.95k arXiv Page

Submitted by

xw-eric

The Unreasonable Effectiveness of Scaling Agents for Computer Use

Behavior Best-of-N (bBoN) improves the reliability and success rates of computer-use agents by generating and selecting among multiple rollouts using behavior narratives, achieving state-of-the-art performance on OSWorld and strong generalization to different operating systems.

Simular · Published on Oct 2, 2025

24

Submitted by

xw-eric

The Unreasonable Effectiveness of Scaling Agents for Computer Use

Behavior Best-of-N (bBoN) improves the reliability and success rates of computer-use agents by generating and selecting among multiple rollouts using behavior narratives, achieving state-of-the-art performance on OSWorld and strong generalization to different operating systems.

Simular · Oct 2, 2025

24

Submitted by

Jeff-Wang

GigaBrain-0: A World Model-Powered Vision-Language-Action Model

GigaBrain-0, a VLA foundation model, uses world model-generated data to enhance cross-task generalization and policy robustness, improving real-world performance on complex manipulation tasks.

GigaAI · Published on Oct 22, 2025

49

GitHub 867 arXiv Page

Submitted by

Jeff-Wang

GigaBrain-0: A World Model-Powered Vision-Language-Action Model

GigaBrain-0, a VLA foundation model, uses world model-generated data to enhance cross-task generalization and policy robustness, improving real-world performance on complex manipulation tasks.

GigaAI · Oct 22, 2025

49

GitHub 867 arXiv Page

Multi-module GRPO: Composing Policy Gradients and Prompt Optimization for Language Model Programs

mmGRPO, a multi-module extension of GRPO, enhances accuracy in modular AI systems by optimizing LM calls and prompts across various tasks.

13 authors

· Published on Aug 6, 2025

GitHub 30.9k arXiv Page

Multi-module GRPO: Composing Policy Gradients and Prompt Optimization for Language Model Programs

mmGRPO, a multi-module extension of GRPO, enhances accuracy in modular AI systems by optimizing LM calls and prompts across various tasks.

13 authors

· Aug 6, 2025

GitHub 30.9k arXiv Page

Submitted by

yulunliu

Generative Refocusing: Flexible Defocus Control from a Single Image

Generative Refocusing uses semi-supervised learning with DeblurNet and BokehNet to achieve high-quality single-image refocusing with controllable bokeh and text-guided adjustments.

3 authors

· Published on Dec 18, 2025

GitHub 64 arXiv Page

Submitted by

yulunliu

Generative Refocusing: Flexible Defocus Control from a Single Image

Generative Refocusing uses semi-supervised learning with DeblurNet and BokehNet to achieve high-quality single-image refocusing with controllable bokeh and text-guided adjustments.

3 authors

· Dec 18, 2025

GitHub 64 arXiv Page

PDFMathTranslate: Scientific Document Translation Preserving Layouts

PDFMathTranslate is an open-source software that translates scientific documents while maintaining layout integrity, utilizing advancements in large language models and layout detection.

4 authors

· Published on Jul 2, 2025

-

GitHub 30.7k arXiv Page

PDFMathTranslate: Scientific Document Translation Preserving Layouts

PDFMathTranslate is an open-source software that translates scientific documents while maintaining layout integrity, utilizing advancements in large language models and layout detection.

4 authors

· Jul 2, 2025

-

GitHub 30.7k arXiv Page

Submitted by

mervenoyan

RF-DETR: Neural Architecture Search for Real-Time Detection Transformers

RF-DETR, a light-weight detection transformer, uses weight-sharing NAS to optimize accuracy and latency for real-time detection across diverse datasets.

Roboflow · Published on Nov 12, 2025

GitHub 4.87k arXiv Page

Submitted by

mervenoyan

RF-DETR: Neural Architecture Search for Real-Time Detection Transformers

RF-DETR, a light-weight detection transformer, uses weight-sharing NAS to optimize accuracy and latency for real-time detection across diverse datasets.

Roboflow · Nov 12, 2025

GitHub 4.87k arXiv Page

Submitted by

zhongwenxu

Single-stream Policy Optimization

Single-stream Policy Optimization (SPO) improves policy-gradient training for Large Language Models by eliminating group-based issues and providing a stable, low-variance learning signal, leading to better performance and efficiency.

Tencent · Published on Sep 16, 2025

GitHub 17.7k arXiv Page

Submitted by

zhongwenxu

Single-stream Policy Optimization

Single-stream Policy Optimization (SPO) improves policy-gradient training for Large Language Models by eliminating group-based issues and providing a stable, low-variance learning signal, leading to better performance and efficiency.

Tencent · Sep 16, 2025

GitHub 17.7k arXiv Page

Submitted by

Owen777

LucidFlux: Caption-Free Universal Image Restoration via a Large-Scale Diffusion Transformer

LucidFlux, a caption-free UIR framework using a diffusion transformer, achieves robust image restoration through adaptive conditioning and SigLIP features without text prompts.

W2GenAI Lab · Published on Sep 26, 2025

21

GitHub 994 arXiv Page

Submitted by

Owen777

LucidFlux: Caption-Free Universal Image Restoration via a Large-Scale Diffusion Transformer

LucidFlux, a caption-free UIR framework using a diffusion transformer, achieves robust image restoration through adaptive conditioning and SigLIP features without text prompts.

W2GenAI Lab · Sep 26, 2025

21

GitHub 994 arXiv Page

Submitted by

Jiasheng1110

Few-Step Distillation for Text-to-Image Generation: A Practical Guide

A systematic study adapts diffusion distillation techniques to text-to-image generation, providing guidelines for successful implementation and deployment.

DAMO Academy · Published on Dec 15, 2025

GitHub 211 arXiv Page

Submitted by

Jiasheng1110

Few-Step Distillation for Text-to-Image Generation: A Practical Guide

A systematic study adapts diffusion distillation techniques to text-to-image generation, providing guidelines for successful implementation and deployment.

DAMO Academy · Dec 15, 2025

GitHub 211 arXiv Page

Submitted by

taesiri

Fara-7B: An Efficient Agentic Model for Computer Use

FaraGen creates synthetic datasets for computer use agents, enabling the training of efficient and high-performing models like Fara-7B on diverse web tasks, outperforming larger models on benchmarks.

Microsoft · Published on Nov 24, 2025

12

GitHub 3.31k arXiv Page

Submitted by

taesiri

Fara-7B: An Efficient Agentic Model for Computer Use

FaraGen creates synthetic datasets for computer use agents, enabling the training of efficient and high-performing models like Fara-7B on diverse web tasks, outperforming larger models on benchmarks.

Microsoft · Nov 24, 2025

12

GitHub 3.31k arXiv Page

Submitted by

taesiri

SAM 3D: 3Dfy Anything in Images

SAM 3D is a generative model that reconstructs 3D objects from single images using a multi-stage training framework that includes synthetic pretraining and real-world alignment, achieving high performance in human preference tests.

AI at Meta · Published on Nov 20, 2025

109

GitHub 5.02k arXiv Page

Submitted by

taesiri

SAM 3D: 3Dfy Anything in Images

SAM 3D is a generative model that reconstructs 3D objects from single images using a multi-stage training framework that includes synthetic pretraining and real-world alignment, achieving high performance in human preference tests.

AI at Meta · Nov 20, 2025

109

GitHub 5.02k arXiv Page

Submitted by

wenbowen

Fast-FoundationStereo: Real-Time Zero-Shot Stereo Matching

Fast-FoundationStereo achieves real-time zero-shot stereo generalization by combining knowledge distillation, blockwise neural architecture search, and structured pruning.

NVIDIA · Published on Dec 11, 2025

4

GitHub 264 arXiv Page

Submitted by

wenbowen

Fast-FoundationStereo: Real-Time Zero-Shot Stereo Matching

Fast-FoundationStereo achieves real-time zero-shot stereo generalization by combining knowledge distillation, blockwise neural architecture search, and structured pruning.

NVIDIA · Dec 11, 2025