📊 Full opportunity report: AMÁLIA · The Three Hard Questions. on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Portugal launched AMÁLIA, a €5.5M European Portuguese large language model, which is now operational. Experts are raising three key questions about its openness, native-language data, and objectives, highlighting broader issues in European sovereign LLM efforts.
Portugal’s €5.5 million investment in the AMÁLIA large language model has resulted in a functional, publicly accessible system, but experts are raising three fundamental questions about its openness, native-language data, and strategic goals, which could influence future European sovereign-LLM initiatives.
AMÁLIA, a consortium project involving approximately 60 researchers from Portugal’s leading research institutions, was announced in December 2024. The model, based on a continuation of the EuroLLM multilingual foundation, was completed on September 30, 2025, and is currently accessible via the FCT’s IAedu platform to 450,000 academic users. It handles Portuguese text with knowledge up to the end of 2023, with a final version expected by June 2026.
The technical approach involves building on an existing multilingual model rather than training from scratch, contrasting with Italy’s Minerva, which trained from zero on Italian and English data. The training pipeline included 107 billion tokens, with 5.8 billion tokens from Portugal’s web archive Arquivo.pt, representing roughly 5.5% of the mixture. Supervised fine-tuning involved 17-18% Portuguese data, with no separate native-language pre-training emphasis.
Preliminary benchmarks show AMÁLIA outperforming previous open models on European Portuguese tasks and beating Qwen 3-8B on most benchmarks, although it still trails on ALBA, the team’s primary Portuguese benchmark. The project remains a work in progress, with the final version expected in mid-2026.
AMÁLIA
The three hard
questions.
Portugal spent €5.5M to build a European Portuguese LLM. The base version is operational, the benchmarks beat Qwen 3-8B on most pt-PT tasks. So why are the most important questions still unanswered?
Last month, Duarte O.Carmo published the sharpest public analysis of AMÁLIA — Portugal’s state-funded European Portuguese large language model. He prefaces his critique with the necessary diplomatic apparatus before doing what almost nobody else in the European-sovereign-LLM discourse has been willing to do publicly: asking hard questions about whether the work, as released, actually does what it set out to do. This piece is a structural extension of his analysis. The AMÁLIA case study exposes three hard questions every national LLM effort needs to answer publicly — and the broader European sovereign-LLM movement has been operating without explicit answers to any of them.
Three questions every national LLM effort needs to answer publicly.
Duarte O.Carmo’s framing maps cleanly onto the structural argument. Each question lands specifically in AMÁLIA — and the broader European sovereign-LLM movement has been operating without explicit answers to any of them.
The three questions form a structural feedback loop. Q3 (optimization target) determines Q2 (data volume needed) which conditions Q1 (openness sufficient for community contribution). The European sovereign-LLM movement collectively benefits from these questions becoming standard methodology disclosure, not exceptional critique.

Mastering LM Studio to Create AI Agents Locally: Master the Art of Local AI Development with LM Studio: A Comprehensive Guide to Building, Optimizing, and Integrating AI Agents
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
107 billion tokens. 5.8 billion clearly pt-PT.
The structurally tractable question with a structurally surprising answer. For a model whose entire stated purpose is European Portuguese prioritization, the native-language share of extended pre-training is 5.5%. The implications cascade into every other question.

Portuguese Flash Cards – Learn Portuguese Language Vocabulary Words and Phrases – Basic Language for Beginners – Gift for Travelers, Kids, and Adults by Travelflips
PORTUGUESE FLASH CARDS – Basic Portuguese words and phrases for beginners and travelers
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Olmo standard. AMÁLIA’s current state.
Allen Institute for AI’s Olmo project defines what “fully open” operationally requires. Olmo doesn’t lead frontier benchmarks. That’s not the point. The point is to be the structural reference for openness. AMÁLIA’s “fully open source” claim should track to the operational standard.

AI Engineering: Building Applications with Foundation Models
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Four strategic positions. AMÁLIA between two and three.
Approximately €100M+ in publicly disclosed European sovereign-LLM funding across the major initiatives. The structural question every project faces: what is the actual competitive position you’re staking? Four options — none mutually exclusive — but each requiring different commitments.

Operating Large Language Models Benchmarking, Deployment, RAG, and Prompt Design (Modern AI Systems Book 5)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Three standards. For AMÁLIA and the movement.
The structural critique generalizes beyond AMÁLIA. Italy, France, Germany, Switzerland, the OpenEuroLLM consortium, and every subsequent national project benefit from public discourse holding national LLM efforts to operational standards on openness, data accounting, and strategic positioning.
The European sovereign-AI agenda is a serious strategic project that deserves serious public discourse. O.Carmo’s analysis is what serious public discourse looks like. Appropriately diplomatic. Structurally rigorous. Willing to ask the hard questions in public when the public investment justifies it. More of this is needed — across every European sovereign-LLM project, not just AMÁLIA.
Implications for European Sovereign-Language AI Strategies
The development of AMÁLIA highlights critical issues facing European efforts to develop national language models. The questions of how open these models truly are, how much native-language data is sufficient, and what their primary objectives should be are central to shaping policy and research directions. The Portuguese case exemplifies how national investments can serve as a testing ground for broader strategic debates, influencing future policies across Europe.
These questions matter because they determine the transparency, cultural relevance, and strategic priorities of European AI initiatives. Addressing them openly can foster more accountable, effective models that serve national interests while contributing to the continent’s technological sovereignty.
European Sovereign LLM Initiatives and Strategic Challenges
Across Europe, multiple countries and consortia—such as Italy’s Minerva, Germany’s Aleph Alpha, France’s Mistral, and others—are developing sovereign language models. These efforts share common structural challenges: defining what it means for a model to be ‘fully open,’ determining the necessary amount of native-language data, and establishing clear strategic goals. The public discourse has often focused on individual model capabilities, but experts argue that understanding the broader structural patterns is crucial for meaningful progress.
Portugal’s AMÁLIA is the first major project to publicly address these questions at a national level, with a transparent investment and open deployment. Its progress provides insights into the technical and strategic trade-offs involved in building European-language large models, highlighting the importance of these overarching questions for the continent’s AI sovereignty.
“AMÁLIA is an impressive piece of work. But the questions about its openness, native data, and goals are essential to understanding its true impact.”
— Duarte O.Carmo
Unanswered Questions About AMÁLIA’s Openness and Strategy
While AMÁLIA is operational and benchmarks are promising, it remains unclear how open the model truly is, especially regarding access, licensing, and data transparency. Additionally, the strategic objectives—whether the focus is on performance, cultural relevance, or policy sovereignty—are still being clarified. The final version’s capabilities and strategic alignment will become clearer only as further benchmarks and policy discussions unfold.
Next Milestones and Ongoing Evaluation of AMÁLIA
The immediate next step is the release of the final version of AMÁLIA in June 2026, which will provide a clearer picture of its capabilities and strategic focus. Over the coming months, researchers and policymakers will scrutinize its openness, data transparency, and alignment with European AI sovereignty goals. Additionally, broader discussions across Europe are expected to address the structural questions highlighted by the Portuguese case, influencing future investments and development strategies.
Key Questions
What makes AMÁLIA different from other European language models?
AMÁLIA is based on a continuation of an existing multilingual foundation, rather than training from scratch, and is publicly funded with a focus on Portuguese language tasks. Its development emphasizes transparency and strategic questions relevant to European sovereignty efforts.
What are the main concerns about AMÁLIA’s openness?
Experts are questioning how open the model truly is, including access restrictions, licensing, and the transparency of its training data, especially native Portuguese sources.
Why do the three questions matter for European AI development?
They determine the transparency, cultural relevance, and strategic focus of models, impacting policy, innovation, and sovereignty across Europe.
When will we see the final version of AMÁLIA?
The final version is expected to be released in June 2026, after which its capabilities and strategic alignment will be better understood.
How does AMÁLIA influence broader European AI efforts?
It serves as a case study highlighting key structural questions that other national models must address, shaping the continent’s AI sovereignty policies.
Source: ThorstenMeyerAI.com