







Turn raw images, videos, audio and 3D data into model-ready datasets with CVAT. Use AI-assisted annotation, quality control, analytics, collaboration tools, APIs, and expert labeling services.
On-device intelligence for every product
We're building a frontier lab for on-device AI. Small, specialized models for audio, vision, and text, faster than the cloud and free.

DeepL AI Platform: Translation, Voice & API
Explore our AI suite and get more done: Translate speech, text, and media, or integrate the DeepL API.
The New SDLC With Vibe Coding
Discover what actually works in AI. Join millions of builders, researchers, and labs evaluating agents, models, and frontier technology through crowdsourced benchmarks, competitions, and hackathons.

HAIM: Human-AI Music Datasets for AI Music Production Tracking Benchmark
As generative platforms such as Suno and Udio reach human-grade audio quality, the scope of AI's utility has expanded across the entire music production workflow. Beyond simple track generation, these advancements have catalyzed the adoption of AI-driven methodologies in diverse forms. These include vocal synthesis, arrangement, and professional mastering. However, current detection research remains largely confined to a binary `AI-or-human' paradigm. It fails to reflect the realities of contemporary music production workflows. In real-world production, AI tools are increasingly used to refine or master human-produced tracks, and human engineers likewise post-process AI-generated material to ensure professional quality. Moreover, users often employ adversarial tactics to bypass AI detectors, such as applying human mastering to AI-generated tracks. This creates a grey area that a simple binary classification fails to capture. In this paper, we define and investigate ``AI Music Tracking'': the challenge of identifying specific AI integration across the multifaceted spectrum of music production. To this end, we introduce HAIM, a dataset with diverse labels for stages of music production. It is designed to isolate stages of AI intervention, including hybrid production and agent-level tracking. Our evaluation of state-of-the-art detectors reveals systemic flaws. By releasing HAIM, we propose a new benchmark that shifts the field beyond binary classification toward a granular, structured evaluation of AI music.

Leanstral: Open-Source foundation for trustworthy vibe-coding | Mistral AI
The most powerful AI platform for enterprises. Customize, fine-tune, and deploy AI assistants, autonomous agents, and multimodal AI with open models.
Muse Platform
Use Meta AI assistant to get things done, create AI-generated images for free, and get answers to any of your questions.
Realtime and audio | OpenAI API
Learn which realtime and audio guide to use for each speech application.

Towards Expert-level Medical AI for Real-time Video Consultations
Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through non-verbal cues. While text-based AI has shown promise, it discards essential perceptual dimensions and limits patients who cannot articulate symptoms in writing. Early efforts to extend medical AI to audio-visual interaction have demonstrated feasibility but not reached clinician-level performance. Here, we provide the first demonstration of expert-level AI in real-time clinical video consultations using AMIE (Articulate Medical Intelligence Explorer) in a video configuration. AMIE (Video) is a Gemini-based multi-agent system integrating low-latency dialogue, clinical reasoning, and real-time audio-visual perception. To guide development, we established a taxonomy and automated evaluations for clinical audio-visual cues in telehealth settings. In a randomized Objective Structured Clinical Examination (OSCE) study with 30 primary care physicians (PCPs), 15 patient actors and 100 clinical scenarios, we compared AMIE (Video), its text-only counterpart AMIE (Text), and PCPs consulting via video. Clinical evaluators rated AMIE (Video) on par or better than PCPs in history-taking, diagnosis, management, and physical observation and examination. Patient actors preferred AMIE's approach to assessing and explaining conditions, while PCPs were preferred for rapport and partnership building. In modality ablation, patient actors preferred AMIE (Video)'s interface over text chat for communicative effectiveness, convenience, and feeling understood. Limitations remain in fine anatomical precision, subtle affective nuances, and high-frequency movements. While further research is needed before real-world translation, these results mark an important milestone toward AI systems capable of augmenting care across the sensory complexity of clinical practice.

Find Open Datasets for AI and Research | Kaggle
Browse and download hundreds of thousands of open datasets for AI research, model training, and analysis. Join a community of millions of researchers, developers, and builders to share and collaborate on Kaggle.

Build Personal AI Agents on Windows PCs with New Tools from Microsoft and NVIDIA | NVIDIA Technical Blog
AI agents are changing how you interact with your PC. Creators, developers, and AI enthusiasts are already using these agents extensively to assist with day-to-day tasks such as coding, video editing…

Pipecat Cloud: Enterprise Voice Agents Built On Open Source - Kwindla Hultman Kramer, Daily
Smart Glasses for Disabled People
This paper introduces an innovative assistive technology aimed at improving the quality of life for individuals with visual impairments. This project presents a system that combines the capabilities You Only Look Once [6] (YOLO) deep learning algorithm with the ESP32 [7] microcontroller to create Object Detection Glasses. These glasses employ real-time object detection to provide users with auditory feedback about their surroundings, enabling them to navigate and interact with their environment more independently. The first task of the glasses is to take pictures of the object and store it as a snap. The second task is that they compare the captured image to the existing image. In order to convert the text into speech, it used Text to Speech technology (TTS). The picture will be taken by ESP 32 with the perfect size the image will be displayed and then by using the speaker the object will be detected. The integration of YOLO [6] and ESP32 [7] offers a cost-effective and efficient solution for enhancing accessibility and inclusivity, promoting greater autonomy and safety for individuals with visual disabilities.
Cryptography may offer a solution to the massive AI-labeling problem
An internet protocol called C2PA adds a “nutrition label” to images, video, and audio.

Detail - Argmax
Detail, Apple's pick for iPad App of the Year 2025, leverages Argmax SDK to build their flagship AI features such as text-based video editing and automatic speaker switching using Argmax SDK, migrating from cloud APIs. - Dec 09, 2025

Evoto - Best AI Photo Editing Software for Professional Photographers
Cull. Retouch. Batch edit. Deliver. Evoto is the all-in-one AI workflow tool trusted by 1M+ photographers. Try free for 7 days.
