Onderzoek RSS

Open Source 🇺🇸

Gigatoken: BPE Tokenizer in Rust Encodes Text at 24.53 GB/s, Up to 989x Faster Than HuggingFace Tokenizers

Stanford PhD student Marcel Roed released Gigatoken, a BPE tokenizer in Rust with Python bindings, achieving speeds of up to 24.53 GB/s on a 144-core AMD EPYC, 989x faster than HuggingFace Tokenizers and 681x faster than tiktoken. The speedup comes from a handwritten SWAR pre-tokenizer and pre-token caching.

OpenAIOpenAI Hugging FaceHugging Face Google/DeepMindGoogle/DeepMind MetaMeta Alibaba/QwenAlibaba/Qwen DeepSeekDeepSeek Moonshot AIMoonshot AI NVIDIANVIDIA MicrosoftMicrosoft Allen Institute for AIAllen Institute for AI Answer.AIAnswer.AI MistralMistral
MarkTechPost24.07 · 01:03
Onderzoek 🇨🇳

Why Are Robots Stuck in Demos? New Company iFlytek Answers: Lack of 'Self-Awareness'

Yao Fang Intelligence, a new company founded by iFlytek, explains that the main problem with robots is not speed but the lack of 'self-awareness' (本体认知). They propose their own architecture iFLYTEK-Embodied-Omni, which integrates VLM, VGM, and AGM to overcome the limitations of VLA and compensate for the shortage of real data. The architecture has already demonstrated high performance on benchmarks including LIBERO-Plus and RoboTwin 2.0.

科大讯飞科大讯飞 NVIDIANVIDIA
QbitAI 量子位24.07 · 01:02
Vers nieuws