🐦 Twitter Post Details

Viewing enriched Twitter post

@SakanaAILabs

We’re excited to introduce KAME: Tandem Architecture for Enhancing Knowledge in Real-Time Speech-to-Speech Conversational AI, accepted at #ICASSP2026! 🐢 Blog https://t.co/arVz1TGpJJ Paper https://t.co/0EwpyRXeCs Can a speech AI think deeply without pausing to process? In real conversation, we don’t wait until we’ve fully worked out what we want to say—we start talking, and our thoughts catch up as the sentence unfolds. Fast speech-to-speech models achieve this, but their reasoning tends to stay shallow. Cascaded pipelines that route through a knowledgeable LLM are smarter, but the added latency breaks the flow—they fall back to "think, then speak." In our new paper, we propose a way to break this trade-off. We call it KAME (Turtle in Japanese). A speech-to-speech model handles the fast response loop and starts replying immediately. In parallel, a backend LLM runs asynchronously, generating response candidates that are continuously injected as "oracle" signals in real time. This shifts the AI paradigm from "think, then speak" to "speak while thinking." The backend LLM is completely swappable. You can plug in GPT-4.1, Claude Opus, or Gemini 2.5 Flash depending on the task without changing the frontend. In our experiments, Claude tended to score higher on reasoning, while GPT did better on humanities questions. Try the model yourself here: https://t.co/uDA0nvvjhS

Media 2

📊 Media Metadata

{
  "media": [
    {
      "url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2049544945233764755/media_0.mp4",
      "media_url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2049544945233764755/media_0.mp4",
      "type": "video",
      "filename": "media_0.mp4"
    },
    {
      "url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2049544945233764755/media_2.jpg",
      "media_url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2049544945233764755/media_2.jpg",
      "type": "photo",
      "filename": "media_2.jpg"
    }
  ],
  "processed_at": "2026-05-11T22:36:59.141185",
  "pipeline_version": "2.0"
}

🔧 Raw API Response

{
  "type": "tweet",
  "id": "2049544945233764755",
  "url": "https://x.com/SakanaAILabs/status/2049544945233764755",
  "twitterUrl": "https://twitter.com/SakanaAILabs/status/2049544945233764755",
  "text": "We’re excited to introduce KAME: Tandem Architecture for Enhancing Knowledge in Real-Time Speech-to-Speech Conversational AI, accepted at #ICASSP2026! 🐢\n\nBlog https://t.co/arVz1TGpJJ\nPaper https://t.co/0EwpyRXeCs\n\nCan a speech AI think deeply without pausing to process?\n\nIn real conversation, we don’t wait until we’ve fully worked out what we want to say—we start talking, and our thoughts catch up as the sentence unfolds.\n\nFast speech-to-speech models achieve this, but their reasoning tends to stay shallow. Cascaded pipelines that route through a knowledgeable LLM are smarter, but the added latency breaks the flow—they fall back to \"think, then speak.\"\n\nIn our new paper, we propose a way to break this trade-off. We call it KAME (Turtle in Japanese).\n\nA speech-to-speech model handles the fast response loop and starts replying immediately. In parallel, a backend LLM runs asynchronously, generating response candidates that are continuously injected as \"oracle\" signals in real time.\n\nThis shifts the AI paradigm from \"think, then speak\" to \"speak while thinking.\"\n\nThe backend LLM is completely swappable. You can plug in GPT-4.1, Claude Opus, or Gemini 2.5 Flash depending on the task without changing the frontend. In our experiments, Claude tended to score higher on reasoning, while GPT did better on humanities questions.\n\nTry the model yourself here: https://t.co/uDA0nvvjhS",
  "source": "Twitter for iPhone",
  "retweetCount": 129,
  "replyCount": 12,
  "likeCount": 644,
  "quoteCount": 31,
  "viewCount": 268654,
  "createdAt": "Wed Apr 29 17:42:48 +0000 2026",
  "lang": "en",
  "bookmarkCount": 404,
  "isReply": false,
  "inReplyToId": null,
  "conversationId": "2049544945233764755",
  "displayTextRange": [
    0,
    279
  ],
  "inReplyToUserId": null,
  "inReplyToUsername": null,
  "author": {
    "type": "user",
    "userName": "SakanaAILabs",
    "url": "https://x.com/SakanaAILabs",
    "twitterUrl": "https://twitter.com/SakanaAILabs",
    "id": "218811492",
    "name": "Sakana AI",
    "isVerified": false,
    "isBlueVerified": true,
    "verifiedType": "Business",
    "profilePicture": "https://pbs.twimg.com/profile_images/1885939209388929024/dtnrOdGp_normal.jpg",
    "coverPicture": "https://pbs.twimg.com/profile_banners/218811492/1686643464",
    "description": "",
    "location": "Tokyo, Japan",
    "followers": 70759,
    "following": 0,
    "status": "",
    "canDm": false,
    "canMediaTag": true,
    "createdAt": "Tue Nov 23 10:20:07 +0000 2010",
    "entities": {
      "description": {
        "urls": []
      },
      "url": {}
    },
    "fastFollowersCount": 0,
    "favouritesCount": 2,
    "hasCustomTimelines": true,
    "isTranslator": false,
    "mediaCount": 380,
    "statusesCount": 1086,
    "withheldInCountries": [],
    "affiliatesHighlightedLabel": {},
    "possiblySensitive": false,
    "pinnedTweetIds": [
      "2036840833690071450"
    ],
    "profile_bio": {
      "description": "Sakana AI is an AI R&D company based in Tokyo. We develop AI solutions for Japan’s needs, and democratize AI in Japan. Try Sakana Chat: https://t.co/1m2lSgnfB2",
      "entities": {
        "description": {
          "urls": [
            {
              "display_url": "sakana.ai",
              "expanded_url": "https://sakana.ai/",
              "indices": [
                136,
                159
              ],
              "url": "https://t.co/1m2lSgnfB2"
            }
          ]
        },
        "url": {
          "urls": [
            {
              "display_url": "sakana.ai/careers",
              "expanded_url": "https://sakana.ai/careers",
              "indices": [
                0,
                23
              ],
              "url": "https://t.co/1q07mb3TzE"
            }
          ]
        }
      }
    },
    "isAutomated": false,
    "automatedBy": null
  },
  "extendedEntities": {
    "media": [
      {
        "additional_media_info": {
          "monetizable": true
        },
        "display_url": "pic.twitter.com/Ut0ypkjJWx",
        "expanded_url": "https://twitter.com/SakanaAILabs/status/2049544945233764755/video/1",
        "ext_media_availability": {
          "status": "Available"
        },
        "id_str": "2049544868536684544",
        "indices": [
          280,
          303
        ],
        "media_key": "13_2049544868536684544",
        "media_results": {
          "id": "QXBpTWVkaWFSZXN1bHRzOgwABAoAARxxcjLwWpAAAAA=",
          "result": {
            "__typename": "ApiMedia",
            "id": "QXBpTWVkaWE6DAAECgABHHFyMvBakAAAAA==",
            "media_key": "13_2049544868536684544"
          }
        },
        "media_url_https": "https://pbs.twimg.com/amplify_video_thumb/2049544868536684544/img/svS94mMh_RNGYwu_.jpg",
        "original_info": {
          "focus_rects": [],
          "height": 1080,
          "width": 1920
        },
        "sizes": {
          "large": {
            "h": 1080,
            "w": 1920
          }
        },
        "type": "video",
        "url": "https://t.co/Ut0ypkjJWx",
        "video_info": {
          "aspect_ratio": [
            16,
            9
          ],
          "duration_millis": 20560,
          "variants": [
            {
              "content_type": "application/x-mpegURL",
              "url": "https://video.twimg.com/amplify_video/2049544868536684544/pl/ladom-knePHGwfPl.m3u8?tag=21&v=874"
            },
            {
              "bitrate": 256000,
              "content_type": "video/mp4",
              "url": "https://video.twimg.com/amplify_video/2049544868536684544/vid/avc1/480x270/wNbLwmvKH5sh4ur9.mp4?tag=21"
            },
            {
              "bitrate": 832000,
              "content_type": "video/mp4",
              "url": "https://video.twimg.com/amplify_video/2049544868536684544/vid/avc1/640x360/r5XSEWbPyrtfgABU.mp4?tag=21"
            },
            {
              "bitrate": 2176000,
              "content_type": "video/mp4",
              "url": "https://video.twimg.com/amplify_video/2049544868536684544/vid/avc1/1280x720/FP_PhlNSWSpbb6UC.mp4?tag=21"
            },
            {
              "bitrate": 10368000,
              "content_type": "video/mp4",
              "url": "https://video.twimg.com/amplify_video/2049544868536684544/vid/avc1/1920x1080/df6Bzk6VDP69Hwbb.mp4?tag=21"
            }
          ]
        }
      }
    ]
  },
  "card": null,
  "place": {},
  "entities": {
    "hashtags": [
      {
        "indices": [
          138,
          149
        ],
        "text": "ICASSP2026"
      }
    ],
    "symbols": [],
    "timestamps": [],
    "urls": [
      {
        "display_url": "pub.sakana.ai/kame/",
        "expanded_url": "https://pub.sakana.ai/kame/",
        "indices": [
          159,
          182
        ],
        "url": "https://t.co/arVz1TGpJJ"
      },
      {
        "display_url": "arxiv.org/abs/2510.02327",
        "expanded_url": "https://arxiv.org/abs/2510.02327",
        "indices": [
          189,
          212
        ],
        "url": "https://t.co/0EwpyRXeCs"
      },
      {
        "display_url": "huggingface.co/SakanaAI/kame",
        "expanded_url": "https://huggingface.co/SakanaAI/kame",
        "indices": [
          1368,
          1391
        ],
        "url": "https://t.co/uDA0nvvjhS"
      }
    ],
    "user_mentions": []
  },
  "quoted_tweet": null,
  "retweeted_tweet": null,
  "isLimitedReply": false,
  "communityInfo": null,
  "article": null
}