Skip to main content
QUICK REVIEW

[论文解读] Sabiá-2: A New Generation of Portuguese Large Language Models

Thales Sales Almeida, Hugo Abonizio|arXiv (Cornell University)|Mar 14, 2024
Language, Linguistics, Cultural Analysis被引用 4
一句话总结

Sabiá-2 是新一代专精于葡萄牙语的大型语言模型,Sabiá-2 Medium 在 64 场考试中击败 GPT-3.5 的 58 场,且在 64 场中 23 场的表现与 GPT-4 相当或更优,同时每 token 的成本仅为后者的 1/10。该模型通过领域专业化实现高性能,且未增加模型规模,在巴西学术与专业考试以及文化背景相关的对话中均表现出色。

ABSTRACT

We introduce Sabiá-2, a family of large language models trained on Portuguese texts. The models are evaluated on a diverse range of exams, including entry-level tests for Brazilian universities, professional certification exams, and graduate-level exams for various disciplines such as accounting, economics, engineering, law and medicine. Our results reveal that our best model so far, Sabiá-2 Medium, matches or surpasses GPT-4's performance in 23 out of 64 exams and outperforms GPT-3.5 in 58 out of 64 exams. Notably, specialization has a significant impact on a model's performance without the need to increase its size, allowing us to offer Sabiá-2 Medium at a price per token that is 10 times cheaper than GPT-4. Finally, we identified that math and coding are key abilities that need improvement.

研究动机与目标

  • 开发新一代专精于巴西学术与专业领域的葡萄牙语大型语言模型。
  • 在无需微调的情况下,仅利用预训练能力,评估模型在多样化、高风险的巴西考试中的表现。
  • 证明通过专业化可提升性能并降低推理成本,且无需增加模型规模。
  • 评估模型在巴西语境下处理文化相关、多轮对话的能力。
  • 识别关键局限性,尤其是数学与编程方面,为未来改进提供方向。

提出的方法

  • 仅使用葡萄牙语文本训练一系列大型语言模型(Sabiá-2),其架构与训练方法未对外公开。
  • 在巴西学术与专业考试的多项选择题子集上评估模型表现,包括 ENEM、BLUEX、ENADE、POSCOMP 等。
  • 使用 GPT-4 作为裁判,在 BRACEval 基准上通过反转提示顺序进行成对比较,以减轻位置偏差,评估开放式对话质量。
  • 评估模型在法律、医学、工程、经济与教育等领域的表现,以衡量其在特定领域中的准确性。
  • 通过领域专业化提升性能,且不增加模型规模,从而实现成本效率。
  • 使用标准化、专家设计的考试,确保公平性与相关性,减少非专家评估带来的偏差。
Figure 1: Results of Sabiá-2, GPT-3.5 Turbo and GPT-4 Turbo on Enade 2022 and 2023 exams, ordered from low to high based on Sabiá-2 performance. Sabiá-2 outperforms GPT-3.5 Turbo on most exams, except Control and Automation Engineering, and Medicine. The lowest accuracies were achieved in domains re
Figure 1: Results of Sabiá-2, GPT-3.5 Turbo and GPT-4 Turbo on Enade 2022 and 2023 exams, ordered from low to high based on Sabiá-2 performance. Sabiá-2 outperforms GPT-3.5 Turbo on most exams, except Control and Automation Engineering, and Medicine. The lowest accuracies were achieved in domains re

实验结果

研究问题

  • RQ1未经微调的专精葡萄牙语大模型是否能在巴西学术与专业考试中超越通用模型如 GPT-3.5 和 GPT-4?
  • RQ2在不增加模型规模或推理成本的前提下,领域专业化能在多大程度上提升模型性能?
  • RQ3与领先专有模型相比,Sabiá-2 在文化背景相关、多轮对话中的表现如何?
  • RQ4该模型的关键局限性是什么,尤其是在数学与编程方面?与当前最先进模型相比有何差异?
  • RQ5AI 裁判(如 GPT-4)能否在开放式、文化敏感的任务中提供可靠且一致的评估?

主要发现

  • Sabiá-2 Medium 在 64 场考试中击败 GPT-3.5 Turbo 的 58 场,展现出在多样化领域中的强大泛化能力。
  • Sabiá-2 Medium 在 64 场考试中 23 场的表现与 GPT-4 相当或更优,尤其在法律、医学与教育领域表现突出。
  • 在基准测试套件中,Sabiá-2 Medium 在 96.9% 的情况下与 GPT-3.5 Turbo 或 Gemini 1.0 Pro 表现相当或更优。
  • 尽管每 token 成本仅为后者的 1/8,Sabiá-2 Medium 在 76.6% 的考试中仍优于 Mistral Large。
  • 在 BRACEval 的多轮对话评估中,Sabiá-2 Medium 在 100% 的有害内容响应中胜出,但在与 Claude 3 Sonnet 和 GPT-3.5 Turbo 的对比中,仅在 15% 的编程问题中胜出。
  • 模型在数学与编程方面面临显著挑战,表明这是未来改进的关键方向。
Figure 2: Performance of Sabiá-2 and other proprietary LLMs on benchmarks of university admission exams: ENEM and BLUEX. The benchmark includes 3 exams.
Figure 2: Performance of Sabiá-2 and other proprietary LLMs on benchmarks of university admission exams: ENEM and BLUEX. The benchmark includes 3 exams.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。