News & Notice

Press Releases

Score Details of the Phase 2 Evaluation of the “Sovereign AI Foundation Model” Project

담당부서
작성자
연락처

The purpose of the “Sovereign AI Foundation Model” project (the Sovereign AI project) is not to make participating teams simply compete for rankings or to eliminate teams. Rather, the project is intended to help Korea’s AI companies grow beyond competition and drive the qualitative growth and expansion of the domestic AI ecosystem.


Accordingly, we have exercised extra caution in determining the scope of disclosure to minimize any unexpected direct or indirect harm that the release of evaluation results could cause to some companies, while also giving balanced consideration to the values of fairness and transparency in the evaluation process and outcomes.


However, on Thursday, August 27, Motif Technologies (“Motif”) requested the disclosure of the detailed evaluation results through a statement. As the other elite teams also agreed to the disclosure, we are releasing the detailed scores for the Phase 2 evaluation of the Sovereign AI project.


In response to the rapid evolution of AI technologies, the Sovereign AI project revises its AI model development goals every six months (a moving target approach). Accordingly, the criteria and methods for each phase of the evaluation are established through consultations with the elite teams.


The criteria and evaluation methods for this Phase 2 were likewise established through several rounds of consultations with the elite teams. The finalized criteria and evaluation methods were formally communicated to the teams before the evaluation and no objections were raised. The evaluation was subsequently conducted on that basis.


As detailed evaluation criteria and key scores were previously disclosed through a press release (August 18) and press briefing materials (August 20), we are now additionally releasing the detailed scores.


The Phase 2 evaluation was conducted based on the following areas: (1) benchmark evaluation (AAII Benchmark Evaluation and NIA Benchmark Evaluation), (2) expert evaluation, and (3) user evaluation (Evaluation by Professional AI Users and Evaluation by the General Citizens).


The AAII (Artificial Analysis Intelligence Index) Benchmark Evaluation


The AAII benchmark evaluation was conducted based on the standards released by Artificial Analysis (AA) at the end of June. AA directly carried out the evaluation and provided the results. As previously announced, the AAII scores were scaled to a maximum of 25 points and reflected in the final scores.


※ Ex: AAII score: 60 → Final score after conversion: 15 (=60 × (25/100))


Based on the overall results of the AAII benchmark evaluation, Motif ranked 1st with 11.9 points, followed by Upstage in 2nd with 9.4 points, SK Telecom (SKT) in 3rd with 8.8 points, and LG AI Research Institute in 4th with 7.8 points.


 


The NIA (National Information Society Agency) Benchmark Evaluation


NIA derived the evaluation results using its benchmark datasets that cover seven areas: mathematics, knowledge, long-text comprehension, safety (including societal safety), reliability, Korean language, and instruction following. The results were then scaled to a maximum of 15 points and reflected in the final scores.


※ Ex: NIA benchmark evaluation score: 60 → Final score after conversion: 9 (=60 × (15/100))


Based on the overall results of NIA benchmark evaluation, SKT ranked 1st with 13.4 points, followed by Upstage in 2nd with 13.3 points, LG AI Research Institute in 3rd with 12.8 points, and Motif in 4th with 12.7 points.


 


The Expert Evaluation


The expert evaluation was conducted by an evaluation committee composed of external experts (three from industry, five from academia, and two from research institutions). Based on the evaluation materials submitted by the elite teams, the committee conducted a written evaluation over approximately five days, alongside Q&A sessions.


The expert evaluation was based on the following criteria: (1) development strategies and technologies (10 points), (2) development outcomes and future plans (10 points), and (3) ripple effects and contribution plans (15 points). In addition, compliance with the minimum requirements for AI sovereignty was also assessed. Each committee member’s score was calculated by adding together the scores for these criteria.


※ For each elite team, the evaluation score was calculated by excluding the highest and lowest scores awarded by the committee members and taking the arithmetic mean of the remaining scores.


Based on the overall results of the expert evaluation, LG AI Research Institute ranked 1st with 29.5 points, followed by SKT in 2nd with 29.3 points, Upstage in 3rd with 29.1 points, and Motif in 4th with 27.1 points.


 


The Professional AI User Evaluation


For the professional AI user evaluation, 50 AI startup CEOs with expertise and experience in AI and more extensive hands-on experience using AI models across various AI development processes than the general public were selected following conflict-of-interest checks. Of these, 49 participated in the evaluation.


The evaluation was conducted based on each team’s AI service website, assessing AI usability and other aspects, with scores ranging from 1 (minimum) to 5 (maximum). The scores from all evaluators were then averaged and scaled to a maximum of 15 points to calculate the final scores.


※ Very poor: 1, Poor: 2, Fair: 3, Good: 4, Excellent: 5

※ Ex: The average score from the scores awarded by 49 professional users: 3.7 → The final score: 11.1 (=3.7×(15/5))


Based on the overall results of the expert evaluation, SKT ranked 1st with 11.6 points, followed by LG AI Research Institute in 2nd with 11.3 points, Upstage in 3rd with 10.8 points, and Motif in 4th with 8.4 points.


 


The General Citizen Evaluation


To ensure that the evaluation was representative of the general citizens, applications were accepted for a sufficient period of time. Quotas were then set by gender and age group based on statistics on Korea’s registered population, and 200 people were randomly selected through the system. Of these, 185 actually participated in the evaluation.


※ All of the above participants were confirmed following checks for conflicts of interest with the four elite teams.


The general citizen evaluation was also conducted as an absolute evaluation for each elite team, with scores ranging from 1 (minimum) to 5 (maximum). The scores from all evaluators were then averaged and scaled to a maximum of 10 points to calculate the final scores.


※ Very poor: 1, Poor: 2, Fair: 3, Good: 4, Excellent: 5

※ Ex: The average score calculated from the scores awarded by 185 general citizens: 3.7 → The final score: 7.4 (=3.7×(10/5))


Based on the overall results of the general citizen evaluation, LG AI Research Institute ranked 1st with 7.6 points, followed by SKT in 2nd with 7.5 points, Upstage in 3rd with 7.3 points, and Motif in 4th with 5.7 points.


 


As a result of aggregating the scores from the preceding evaluations, SKT ranked 1st with 70.6 points, followed by Upstage in 2nd with 69.9 points, LG AI Research Institute in 3rd with 69.0 points, and Motif in 4th with 65.8 points.


 


Even after the disclosure of the detailed scores, we hope that Korean companies will not be judged simply by their scores or rankings, but will instead build on their strengths and capabilities to make broad contributions to the domestic and global AI ecosystem. We also hope that Korea’s AI ecosystem will continue to develop into a dynamic ecosystem where new and competitive AI companies can continually rise to the challenge.



For further information, please contact the Public Relations Division (Phone: +82-44-202-4034, E-mail: msitmedia@korea.kr) of the Ministry of Science and ICT.


Please refer to the attached PDF.

KOGL Korea Open Government License, BY Type 1 : Source Indication The works of the Ministry of Science and ICT can be used under the terms of "KOGL Type 1".
TOP