1
0 Comments

AI Daily Report "May 26 ~ May 27"

1⃣
🎓Google Project Astra acts as an AI tutor:
-Guide students to solve problems independently, emphasizing guided teaching rather than giving direct answers

2⃣
🗣Kyutai launches voice model Unmute:
-Can quickly add voice capabilities to any large text model, modular design

  • Realize intelligent conversation rhythm: can judge whether you have finished speaking and also support user interruption

3⃣
🧠Ali released QwenLong-L1-32B long context reasoning model:

  • The context length supports up to 130,000 tokens, designed specifically for long text tasks
  • Strengthen learning optimization to improve the ability to transfer reasoning from short to long contexts

4⃣
🎙Claude launches voice assistant:
-Support access to personal information sources such as Calendar, Gmail, Google Drive, etc.
-Can conduct online searches and provide answers based on the results, improving the practicality of the smart assistant

5⃣
🌍Claude web search function is now available worldwide:

  • Available to free users worldwide, no additional subscription required

6⃣
💆‍♀Enhancor AI: De-AIing AI-generated images
-Focus on solving the problem of unnatural skin texture in AI images

  • Provides automatic skin texture restoration and detail enhancement

7⃣
🧭Google Project Mariner: Natural Language Scheduling AI Intelligent Agent
-Users can use natural language to schedule multiple agents to perform complex tasks in the browser

8⃣
🎬AI Field Shooting Tutorial: Google Street View + Runway Synthesis of Real Scenes
-Use Google Street View and Runway to "place" people into real-world street scenes

  • Create videos that feel like they were shot on location without having to go out to shoot

9⃣
🧠Claude 4 Best Practices for Prompt Word Engineering:

  • Covers Opus 4 and Sonnet 4 versions.
  • Distilled the 7 golden rules used by Claude 4 to improve the prompt effect.

🔟
🥯BAGEL: ByteDance’s open source multimodal big model is unveiled!

  • Open source alternatives to challenge GPT-4o and Gemini 2.0.
  • Supports image and text input and output, and has image editing, reasoning, and combined modeling capabilities.

11
🧬Three new Gemma model variants released:

  • MedGemma: A medical AI model that can run on a single GPU.
  • SignGemma: An assistive technology model focused on sign language recognition and generation.
  • DolphinGemma: Explore cross-species communication with dolphins and other species.
on May 28, 2025