Wikidata Community Summit @ COSCUP 2026
2026. Aug. 9th (Saturday)
Under the storm of Typhoon Dolphin, COSCUP 2026 commenced, bringing partners from around the world together to celebrate the value of open technology. The Conference for Open Source Coders, Users & Promoters (COSCUP) is one of the biggest open source events in Taiwan. This year, one topic dominated the community: Artificial Intelligence, or Agent in the concurrent stage.
This year, our joint track of Wikidata and OpenStreetMap not only showcased how our respective platforms contribute to the open community, but also explored how open data and open technology can support the rapidly developing world of artificial intelligence.
One highlight was our presentation with our partner from Wikimedia Deutschland, focusing on the Wikidata MCP Server and the Wikidata Embedding Project.
Artificial Intelligence has come to the stage like a storm and quickly become one of the most discussed technologies across every industry. No matter which side of the fence you are on, one thing is the same: we want to know more about it and how it may affect our lives.
However, much of the discussion throughout the internet is filled with empty buzzwords and misinformation from companies with various interests. The problem is not necessarily malicious intent but the fact that people, including the sources they are quoting, simply do not have enough access to the information they need to understand the technology they are discussing.
Just like previous technology βboomsβ, such as cryptocurrency and NFTs, Artificial Intelligence, and now the so-called Agents, is a market phenomenon of attention crowding around something new and potentially profitable, and thus money floods in. However, unlike its predecessors, AI is βrealβ in a sense that it has the potential to reshape, and possibly damage, many systems of trust that we humans have built through centuries of mutual understanding and respect for ourselves and others.
AI Agents are not inherently disruptive or destructive; it is how we use them that is causing all the trouble. If we want the technology to become a tool for a better future, we first need to understand how the tool works.
The Problem with LLMs
A Large Language Model, or LLM, at its simplest, is an algorithm that predicts words based on the material it has been provided. Its strength comes from recognizing patterns in a given language and mimicking the βnatural languageβ according to the probability of words appearing together. But this also reveals its fundamental limitation.
All it captures is the distribution of words appearing in languages, not necessarily the why and how behind them. It can produce an answer that looks and sounds correct without having any understanding of whether the information corresponds to reality or is actually true. Hallucinations, gibberish, and blatantly false answers are ultimately extensions of this limitation.
The man in the box never understands Chinese; it just appears to.
This does not mean we need the LLM to become a genie in a lamp that knows everything. The bottleneck is never the generation of more knowledge. It is discovering, retrieving, and acting upon the right information. To achieve this, there is no reason to reinvent the wheel of synthesizing the generation of natural knowledge; instead, what we need is to have the preexisting systems align properly.
And this is where Wikidata comes into play.
A Beacon for AI Agents
Wikidata is a Knowledge Graph: a machine-readable database where information is connected through controlled and structured relationships. It is designed to work across languages, interfaces, databases, services, and users β human and machine.
More importantly, its information is human-pruned and anchored in reality. It can provide explicit and exhaustive relationships, traceable sources, and continuously updated information. This makes Wikidata particularly valuable to LLMs. Rather than leaving an AI Agent to decide whether a string of text is factual or fictional, it can reference Wikidata for additional information.
LLMs excel at semantics, but that strength is inherently imprecise. A Knowledge Graph provides the controlled structure underneath it. Instead of asking an Agent to know everything, we can let it discover relevant information through the web of knowledge and crawl through information more efficiently than we humans could.
This is where the Wikidata Embedding Project aims to establish its influence.
The project prepares a vector database from a Wikidata dump and an MCP Server that allows AI Agents to access knowledge residing in it through a standardized and LLM-friendly interface. It combines the fuzzy discovery capabilities of AI with the structured and human-pruned knowledge of Wikidata.
The advantage is not only accuracy. LLMs are trained primarily on mainstream data, with limited information from less represented communities. Wikidata can provide access to knowledge beyond an Agentβs original training, including information from smaller communities that would otherwise be extremely difficult for an AI to discover and access.
By equipping an Agent with Wikidata, it is like a sailor on a misty sea finally seeing the light of a beacon. The Agent can still navigate on its own, but it now has a solid reference point so it wonβt get lost as long as it can see the light.
The Current and the Future
Wikidata is still one of the youngest WikiProjects. Though volunteers around the globe are working diligently, there are still gaps in the Knowledge Graph. If we want Wikidata to become the backbone to power the next generation of Artificial Intelligence, we still have a long journey ahead of us.
The emergence of AI Agents brings as many opportunities as it does challenges. For Wikidata Taiwan, our work is becoming increasingly relevant to making sure Taiwanβs audience and data wonβt get left behind in the ever-changing landscape of Artificial Intelligence.
The challenge is not simply to make AI more powerful. It is to make sure that the knowledge it discovers is reliable, traceable, and based on reality, and most importantly, fair.