Уровень 0 · материалов: 8
В кластер включены документы о создании и выпуске конкретных архитектур больших языковых моделей, оптимизированных для русского языка, и исключены материалы, не касающиеся разработки или технических характеристик LLM.
Общие признаки: выпуск LLM, разработка моделей на русском языке, технические бенчмарки и метрики, открытый исходный код
Группа выше: Конкретные модели и ИИ-платформы
Смысл: The main idea is to present GigaChat as a competitive, Russian-centric alternative to ChatGPT, detailing the technical journey from a base ruGPT model to an instruction-tuned assistant while promoting Sber's contributions to the open-source AI community.
Sber introduces GigaChat, a Russian LLM based on ruGPT-3.5 13B, detailing its instruction-tuning process, multimodal capabilities, and plans for future expansion.
Смысл: The main idea is to present GigaChat MAX as a state-of-the-art Russian LLM that matches or exceeds global competitors in specific benchmarks, with a strong emphasis on the synergy between high-quality expert data (STEM), aesthetic response formatting, and technical optimization.
Sber unveils GigaChat MAX, a high-performance LLM featuring superior STEM capabilities, structured 'beautiful' responses, and competitive benchmarks against GPT-4o.
Смысл: The text announces the launch of GigaChat 3 Ultra and Lightning, highlighting the technical achievement of training a massive 702B MoE model from scratch in Russia to provide a high-performance, open-source alternative that natively excels in the Russian language.
SberDevices releases GigaChat 3 Ultra, a 702B parameter open-source MoE model trained from scratch, setting a new benchmark for native Russian language LLMs.
Смысл: The main idea is to showcase the technical evolution of GigaChat, proving through benchmarks that it is now competitive with or superior to ChatGPT in specific metrics and offers expanded context capabilities for handling larger documents.
GigaChat Pro has outperformed ChatGPT (GPT-3.5) in certain MMLU benchmarks, and GigaChat Lite+ has expanded its context window to 32k tokens.
Смысл: The main idea is to announce the release of an improved version of GigaChat, demonstrating its technical superiority over previous iterations through metrics and explaining the underlying technology of LLMs to the reader.
Sber introduces an upgraded GigaChat model with better performance metrics and a deeper explanation of how neural language models evolve from rule-based systems.
Смысл: Sber has democratized access to high-capacity Russian language models by releasing ruGPT-3, showcasing the company's computational power and encouraging the Russian research community to build new AI-driven products.
Sber has released open-source Russian GPT-3 models (Medium and Large) trained on its 'Christofari' supercomputer and a 600 GB dataset to foster AI development in Russia.
Смысл: The text announces the open-source release of ruGPT-3.5, a 13B parameter language model optimized for Russian, to empower the developer community and advance Russian NLP.
Sber has open-sourced ruGPT-3.5, a 13-billion parameter language model optimized for the Russian language, available under the MIT license on Hugging Face.
Смысл: T-Bank has released two high-performing, open-source Russian LLMs (T-Lite and T-Pro) based on Qwen 2.5, utilizing a cost-effective continual pre-training approach to outperform most competitors in Russian language benchmarks.
T-Bank releases T-Lite (7B) and T-Pro (32B), state-of-the-art open-source Russian language models based on Qwen 2.5, available on Hugging Face.