Open Source
AIEZZ tool tag
Browsing AI products tagged “Open Source”, with 29 matching results.
SmartSub – Open-Source All-in-One Desktop Tool for Audio and Video Subtitle ProcessingAI Speech RecognitionSmartSub (妙幕) is an open-source, all-in-one desktop tool for audio and video subtitle processing. It integrates the entire workflow into a single application, using local models such as Whisper, FunASR, and Qwen3-ASR for offline speech-to-text, supporting over 20 translation services and multi-role AI dubbing. It includes a built-in online video downloader compatible with platforms such as Bilibili and YouTube, supports hardware acceleration including NVIDIA CUDA and Apple Core ML, and runs locally across Windows, macOS, and Linux.00
InstructAV2AV – BAAI and Peking University’s Open-Source Joint Audio-Video Editing ModelImage EditingInstructAV2AV – BAAI and Peking University’s Open-Source Joint Audio-Video Editing Model is a joint audio-video editing model released by the Beijing Academy of Artificial Intelligence and Peking University. With just one natural-language instruction, users can edit the video visuals and corresponding audio simultaneously during end-to-end generation, without manual masks or step-by-step processing. The model supports dialogue modification, character replacement, object insertion and removal, while preserving background and ambient sounds and maintaining audio-visual synchronization. Core capabilities include identity-preserving speech editing, audio-video instance replacement, instance insertion, instance removal, and compositional attribute editing. It is suitable for film and television post-production, short-video creation, virtual human operations, advertising ideation, and interactive content production.00
JoyAI-Video-EditAI Video EditingJoyAI-Video-Edit – JD.com's open-source real-time streaming video editing model is a self-developed, open-source real-time streaming video editing model from JD.com. Built on a 16B-parameter autoregressive diffusion architecture, it achieves 30 FPS inference at 720P resolution and supports stable streaming editing for videos of any duration. Users can modify people, scenes, styles, and objects in real time during video playback through natural-language instructions, without waiting for the complete video to be generated. The model surpasses all streaming editing methods on the OpenVE-Bench evaluation, leading across all metrics, while also providing a large-scale data synthesis pathway for embodied intelligence.00
Orchard – Microsoft Research’s Open-Source Agentic AI Modeling FrameworkAI AgentsOrchard – Microsoft Research’s open-source Agentic AI modeling framework. Centered on the Kubernetes-based Orchard Env environment service, the framework supports cross-domain sandbox reuse for data distillation, reinforcement learning rollouts, and evaluation, covering three major scenarios: software engineering, browser navigation, and personal assistants. It provides complete SFT + RL training recipes, along with 107K SWE trajectories and 3,070 GUI multimodal rollout data points, and enables end-to-end training directly in real harnesses such as Codex and OpenClaw. Designed for academic and industrial teams conducting Agent research and development, Orchard offers lower compute and hosting costs while helping smaller models approach frontier performance through cross-harness generalization.00
Alpamayo 2 Super – NVIDIA’s Open-Source Autonomous Driving AI Reasoning ModelAI ProgrammingAlpamayo 2 Super is NVIDIA’s open-source autonomous driving AI reasoning model, built on NVIDIA Cosmos 3 Super Reasoner. It provides 360° environmental perception, advanced driving decision-making, and automatic reasoning label generation. The model ranks first on the LingoQA autonomous driving reasoning benchmark and uses the permissive OpenMDW-1.1 license, supporting fine-tuning, derivatives, and commercial redistribution. Alpamayo 1.5 / 1 are also provided for cloud development and in-vehicle distillation deployment, targeting Robotaxis and autonomous vehicles.00
wigoloAI Agentswigolo is an open-source local search tool that connects to coding agents such as Claude Code and Cursor via MCP. No API key is required, enabling free local search, web fetching, site crawling, structured extraction, ...00
MemHarness – An Agent Memory Reconstruction Framework by Shanghai AI Lab and OthersAI AgentsMemHarness is an LLM Agent memory reconstruction framework released by Shanghai AI Lab in collaboration with Zhejiang University, Fudan University, Shanghai Jiao Tong University, and other universities. It inserts critique and reconstruction steps between retrieval and action, using the current context to retain, rewrite, or discard historical experiences. It uses GRPO for end-to-end training and requires no additional human annotations. The framework supports the complete workflow of memory retrieval, critique and reconstruction, action generation, experience write-back, and pruning, and provides a five-stage decision-making process. Its 7B-parameter model surpasses Gemini-2.5-Pro by 23.1% and 39.7% on ALFWorld and WebShop, respectively, achieving an average success rate of 85.9% in out-of-distribution scenarios. It is suitable for embodied household tasks, autonomous online shopping, intelligent customer service, code-assisted development, and other scenarios.10
ForgeStencil – ForgeStencil: FaceWall Intelligence’s Fully Automated Stencil Research and Deployment SystemAI AgentsForgeStencil – FaceWall Intelligence’s fully automated Stencil research and deployment system is jointly developed by FaceWall Intelligence and the OpenBMB community. It adopts a dual-agent closed-loop architecture consisting of a Kernel Agent and an App Agent. The system autonomously completes the entire process, from discovering optimization strategies and synthesizing high-performance CUDA Kernels to production-grade application integration and validation, with zero human intervention. Within one week, the system completed end-to-end optimization for more than 100 industrial and scientific computing software applications, achieving a median speedup of 1.41×. It covers 8 major industrial and 5 major scientific domains, including oil and gas exploration, electromagnetic simulation, and medical imaging, significantly reducing reliance on HPC experts and shortening development cycles.00
dots.ttsAI Voice Synthesisdots.tts is a 2-billion-parameter fully continuous autoregressive speech synthesis foundation model jointly open-sourced by the Xiaohongshu dots team and the X-LANCE Lab at Shanghai Jiao Tong University. The model directly generates 48kHz audio chunk by chunk in a continuous latent space, achieving SOTA voice similarity and content accuracy on benchmarks such as Seed-TTS-Eval, while supporting zero-shot cloning, streaming output, and low-latency duplex dialogue. ...00
ShieldstralAI Q&AShieldstral – an open-source multimodal content safety classification model released by Mistral AI – is a 3B-parameter open-source multimodal content safety classification model built on Ministral-3B. The model redefines traditional fixed-category content moderation as a binary question-answering task, supporting the real-time definition of moderation policies through natural-language queries and adaptation to different scenarios without retraining. The model achieves an average F1 score of 84.9% on text safety benchmarks and a multimodal safety F1 score of 83.8%, both reaching SOTA performance. It can be deployed on a single 16GB GPU and supports 12 languages. Core features include policy-adaptive moderation, unified multimodal detection, fine-grained violation identification, calibrated safety scoring, and lightweight edge deployment. It is suitable for scenarios such as content risk control on social platforms, AI dialogue safety protection, advertising and marketing compliance, online education content filtering, and enterprise document security auditing.00
PAST-BenchAI AgentsPAST-Bench is a benchmark introduced by Mengdi Wang's team at Princeton University to evaluate the recursive self-improvement capabilities of personal AI agents. By comparing task performance under conditions with and without memory, PAST-Bench deter...00
PixelRAG – Berkeley’s Open-Source Vision-Native RAG FrameworkDevelopment PlatformsPixelRAG is a vision-native RAG framework open-sourced by Berkeley SkyLab/BAIR. It moves beyond the traditional approach of extracting webpages into plain text before retrieval, instead using a browser to render webpages and PDFs into screenshot tiles, with LoRA-fine-tuned...00
LTX-2.5AI Video GenerationLTX-2.5 is LTX's open-source foundational AI video generation model with 22B parameters and open weights available from launch. The model supports native multishot generation, 4K HDR and RAW/EXR professional workflows, synchronized audio generation, and precise video editing capabilities, ...00
HOMIEAI Video GenerationHOMIE is an open-source digital human video generation framework from the Hong Kong University of Science and Technology, built on the Wan2.1-T2V-14B backbone and integrated with the Qwen3-VL multimodal large language model. The framework uniformly handles four elements—digital humans + products + logos + text—...00
Palmier ProAI Video EditingPalmier Pro is a native macOS video editor built for the AI era. Developed from the ground up in Swift, it is open source with free core features. The product embeds generative AI directly into the timeline and supports Seedance, Kling ...00
MindMemOS – Huawei Noah's Ark Lab's Open-Source AI Agent Memory Operating SystemAI AgentsMindMemOS – Huawei Noah's Ark Lab's open-source AI agent memory operating system, designed to provide agents requiring long-term memory management with an open, reusable memory layer. It uses a three-dimensional entity–attribute–time structure to decouple memory from individual agents and turn it into an independent, cross-framework asset. Through mechanisms such as Dreaming for offline organization, Feedback for correction, and Skill Evolution for automated improvement, it continuously optimizes memory. The system provides dual-mode memory writing (MindVanilla and MindSchema) and multi-path Compact Search retrieval, supporting Docker local deployment, cloud APIs, a Python SDK, CLI, and other integration methods. With over 94% structured recall accuracy and a 10% improvement in question-answering accuracy after memory compression, it helps developers share information across sessions and improve task efficiency in scenarios such as personalized assistants, multi-agent collaboration, and complex office automation.00
Nemotron 3.5 LightningAI AgentsNemotron 3.5 Lightning is NVIDIA's open-source 30B-parameter MoE model, optimized for multi-agent systems. Compared with similar models, it delivers 4× faster output and accelerates agent task completion by 30%, while...00
Wan-Animate-2 – Wanxiang Team’s Open-Source Next-Generation Character Animation ModelAnimation ProductionWan-Animate-2 is a new-generation character animation model open-sourced by the Wanxiang team. As a major architectural upgrade to Wan-Animate, it uses an end-to-end dual-branch DiT and can capture motion priors from reference videos without explicit pose skeletons. It supports driving-video-to-character animation, identity and appearance preservation, and text-driven camera-view control. It provides two inference modes: Base for high quality and Distillation for fewer steps, and introduces the Lite variant for real-time streaming character animation. The model is suitable for virtual streamers, short-video creation, digital humans, and other applications, aiming to improve both motion fidelity and identity consistency while reducing generation latency.00
SenseNova U1.5-Lite-Preview – SenseTime's Open-Source Lightweight Multimodal ModelAI Large Language ModelsSenseNova U1.5-Lite-Preview – SenseTime's open-source lightweight multimodal model is based on the NEO-Unify architecture. With 8B-MoT parameters, it natively unifies visual understanding, reasoning, generation, and editing. The model supports native 4K ultra-high-definition image generation, with fine textures, realistic details, and Chinese and English text rendering in complex layouts. Its built-in Prompt Enhance Skill lowers the barrier to creation, enabling precise execution of long natural-language and structured visual instructions, as well as continuous iterative editing. Designed for designers, developers, and content creators, it provides a commercial-grade visual creation tool with low deployment costs and strong instruction-following capabilities.10
IndexTTS-2.5AI Voice SynthesisIndexTTS-2.5 is an industrial-grade zero-shot voice cloning model open-sourced by Bilibili. With only 0.8B parameters, it supports cross-lingual transfer across Chinese, English, Japanese, Spanish, and Arabic.00
MAGI-2-preview – Sand.ai’s Open-Source Multimodal Video Generation ModelAI Video GenerationMAGI-2-preview – Sand.ai’s Open-Source Multimodal Video Generation Model is Sand.ai’s open-source, 100-billion-parameter MoE multimodal video generation model, with 114B total parameters and only 6B activated. The model retains a single-stream architecture that incorporates text, video, and audio into the same Transformer for joint modeling, enabling native multimodal generation. Its core capabilities include text-to-video, image-to-video, and video-to-video generation, along with native audio understanding. Sparse MoE activation significantly reduces inference costs. Designed for AI researchers, developers, and enterprise technical teams, it supports academic research, private deployment, and multimodal film and television previsualization, providing an efficient, open-source, commercially usable video generation solution.00
Qwen-CUA – A Native Computer Use Agent from Alibaba Qwen and OthersAI AgentsQwen-CUA – a native Computer Use Agent developed by Alibaba Qwen and others – is built on a 397B-A17B mixture-of-experts architecture. It perceives interface states solely through screenshots, without relying on the DOM tree or accessibility metadata, and directly outputs keyboard and mouse events to control browsers, desktops, and professional software. It scores 86.2 on the OSWorld-Verified benchmark, while the larger Qwen-CUA-Max reaches 87.6, placing it among the leading open-source CUA systems. With end-to-end vision-action training, it offers cross-platform versatility and high stability. Its technical report and demo are open source, supporting use cases such as automated testing, RPA, accessibility assistance, and control of scientific software.00
SALMONN-2 – General-Purpose Audio Large Language Model Open-Sourced by Tsinghua University and OthersAI ProgrammingSALMONN-2 – General-Purpose Audio Large Language Model Open-Sourced by Tsinghua University and Others is a general-purpose audio large language model built on ByteDance's SPEAR unified self-supervised audio encoder and a multi-layer feature fusion adapter, with Qwen3 as its text backbone. It supports cross-domain tasks including speech recognition, audio captioning, music understanding, emotion recognition, and speaker verification. It is also the first general-purpose ALLM to systematically implement advanced capabilities such as sound event detection, audio deepfake detection, speech quality assessment, and multimodal in-context learning for speech recognition. The model uses a single SPEAR encoder instead of the traditional dual-encoder architecture and fuses multi-layer acoustic features through an MLF adapter to achieve more balanced cross-domain performance. With efficient training on only approximately 18,000 hours of supervised data, SALMONN-2 achieves the best performance among open-source models of comparable size on the three comprehensive benchmarks MMAU-Pro, MMAR, and MMSU, making it suitable for applications such as intelligent meeting assistants, audio content moderation, and speech quality monitoring.00
Muse Glimmer – Meta's Open-Source 30B-Parameter Large Language ModelAI AgentsMuse Glimmer is Meta's open-source 30-billion-parameter large language model, deeply optimized for always-on local Agent workflows. With 4-bit quantization, the model runs smoothly on a single GPU with 24 GB of VRAM or Apple Silicon (M4/M5 Max) with 32 GB of unified memory, with nearly zero performance degradation. Integrated DFlash speculative decoding delivers up to 3.1× faster performance on high-end hardware such as the RTX 5090. The model supports multimodal tasks including real-time conversation, tool calling, code generation, and multi-step reasoning. It is suitable for individual developers, small and medium-sized businesses, and privacy-sensitive industries such as healthcare and finance, providing efficient local AI capabilities with a low VRAM requirement while keeping all data on the device.00
Hunyuan3D-Buffalo 1.0 – Tencent Hunyuan’s 3D Multimodal Framework3D DesignHunyuan3D-Buffalo 1.0 – Tencent Hunyuan’s 3D Multimodal Framework is a unified 3D multimodal framework developed by Tencent Hunyuan. With a shared Hunyuan3D-VLM backbone, it integrates 3D question answering, spatial grounding, text-to-3D generation, instruction-based editing, and part generation into a single workflow. The framework supports understanding 3D model structures through natural language, locating parts, modifying local geometry based on instructions, and extracting semantic-level components. It aims to connect the full cycle of 3D understanding, generation, and editing, providing a composable 3D asset production solution for games, animation, and industrial design. Core capabilities include 3D understanding, text-to-3D generation, instruction-based editing, and part generation, with all four task types sharing the same 3D semantic representation and processing pipeline. It is suitable for game developers, animators, industrial designers, and 3D content creators, lowering the professional barrier through natural-language interaction while improving asset reuse and iteration efficiency.00
ShotcutAI Video EditingShotcut is open-source video editing software based on FFmpeg, completely free with no watermark or time limits. It supports editing hundreds of audio and video formats directly without importing, and provides a multitrack timeline, keyframe animation, rich...00
WorldClaw – A 3D World Generation Framework from Tencent Hunyuan Team3D DesignWorldClaw is an Agent-driven open-world 3D generation framework developed by the Tencent Hunyuan team. With a single natural-language description, the system automatically plans terrain, regions, buildings, objects, and spatial relationships. Using a coarse-to-fine generation strategy, it builds complete 3D scenes in Blender that can be freely explored and edited object by object. The framework is powered by multiple collaborating models, including Claude Opus, GPT-Image-2, SAM3, and Hunyuan3D. It also features visual Agent-based self-checking and repair, automatically correcting contact issues such as floating objects and mesh intersections.00
PyTorchAI Coding ToolsPyTorch is an open-source machine learning framework initiated by the Meta AI research team and now managed by the Linux Foundation. It primarily serves AI researchers, machine learning engineers, data scientists, and students. Users typically use Python to define models and forward computation flows, then rely on tensor computation, automatic differentiation, and neural network modules for training, while using related ecosystem libraries for vision, audio, or text tasks. Dynamic computation graphs facilitate experimentation and debugging, while also supporting the progression of research prototypes toward model deployment. Practical use still requires knowledge of Python, tensor operations, and deep learning fundamentals. Moving from experimental code to production also requires independently handling engineering deployment and resource configuration.057.6
Qwen 3.8-Max – The flagship large model launched by Alibaba Cloud’s Qwen teamAI Large Language ModelsQwen 3.8-Max – the flagship large model launched by Alibaba Cloud’s Qwen team, expanding the Qwen 3.5 architecture to 2.4 trillion parameters and positioned as an ultra-large language model for enterprise and research use. It supports long-horizon autonomous coding, paper reproduction, competition practice, workplace productivity, quantitative strategy development, and autonomous chip design. Through real-world reinforcement learning and a self-evolving feedback loop, the model can execute autonomously for days to tens of days and complete complex tasks without human intervention. Users can call it through APIs compatible with the OpenAI or Anthropic protocols on the Qwen AI platform, or deploy its open-source weights locally for easy integration with agent frameworks such as Claude Code, Codex, and Qoder. As a result, professional users in R&D, legal and compliance, design, and other fields can significantly improve productivity, compressing long-cycle projects into deliveries completed within a single conversation.10All tools with this tag are shown.
