Reading view

There are new articles available, click to refresh the page.

thundersnap 0.01: an undo button for everything

Happy July 4th! For those of us around the world contemplating independence, it's a good day to think about how we came to rely on expensive cloud infrastructure for our fundamental computing needs.

With that in mind, here is my latest toy project: an open source tool that makes replicating, forking, sharing, and running container snapshots fast and easy across cloud and personal devices.

It's fun to play with, especially on bare metal hardware you run at home, or rent from a provider like Hetzner or OVH. Or, because it uses Tailscale, why not all of them in a single mesh?

There's a lot more to say but I don't have time right now. Details are in the README.

I will say this: humans and AI agents both want the same things when they're trying to get work done. Ephemeral containers aren't really it. But how about unlimited disk space, fast CPUs, an undo button, and the ability to move to whatever provider offers the best hardware at the best price? That's more like it.

Go visit thundersnap on github and tell me what you think!

Wikimania 2026: Growing language communities

Are large language models widening the digital divide between majority and minority languages? Or can they be harnessed in preserving language diversity?

On Day 3 of Wikimania Paris, representatives of many different language and cultural communities shared their experiences in working within the Wikimedia movement.

Respect for indigenous communities

Erroneous information about indigenous languages and cultures can alienate potential contributors. During the keynote panel, Michelle Collipal, a member of the indigenous Mapuche community in Chile, shared how she has started to motivate Mapudungùn speakers and cultural authorities to engage with the Mapudungùn Wiki project. 

Moderator Dr Terri Janke underscored the importance of involving indigenous communities, respecting their knowledge as well as their right to self-determination. What this means in practice for wiki spaces is discussed in the white paper “CutureStrong Platforms: Setting the Standard at Wikimedia”.

Automating and improving translation

During “Language Diversity in the Digital Commons”, Pau Giner of the Wikimedia Foundation shared that the MinT (Machine in Translation) was designed from the start to support diverse language models. The platform currently supports 200 languages.

His ideas for the future include making Wikipedia articles more accessible on mobile phones, and to offer a collection of tools to help people create their first wiki, going beyond translation.

During the Q&A, Kepa Sarasola (User:karasola) of University of Basque reflected on the improvement of translation tools available since he started working on Basque Wikipedia in 2011. The tools, he said, had made it possible to grow minority language Wikipedias – while giving them the flexibility to choose which parts of articles to translate and where to start from scratch. The integration of AI tools has accelerated the process significantly.

Collaboration across communities

The need for collaboration across language communities was a recurring theme across multiple sessions. One example of a cross-border network is Linguatec-IA, an EU project for the digitization of languages in communities in the Pyrenees region, including Basque, Catalan, Occitan, and Aragonese.

David Castillo Parra of UNESCO discussed the launch of the New Commons Incubator for indigenous-led capacity building programs. Applications will be accepted for indigenous-led teams through 14 August.

The need for the Wikimedia movement to remain true to its mission – and to aligning on milestones that matter – was an inspiring message from Audrey Tang in “The State of Wikimedia & AI 2026”. “Don’t let it be a race…Wikipedia never tried to win. We tried to make sure there were many winners at any one time.”

As Jimmy Wales told Le Monde, “It’ll be all right. We’ll adapt, change, use AI in our own way. We’ll find a way.”

Poland and Ukraine joined forces in documenting heritage: results of the 2025 Wiki Loves Monuments campaign

For the second time in a row a Ukrainian edition of Wiki Loves Monuments international photo contest had a special category dedicated to Polish heritage in Ukraine — almost 4600 photos by more than 100 authors were submitted. This special category was a joint project of Wikimedia Polska and Wikimedia Ukraine.

A collage of the winning photos of the 2025 Polish Heritage in Ukraine campaign

Mykola Kozlenko (NickK), a member of Wiki Loves Monuments Ukraine organising team, Board member of Wikimedia Ukraine, commented:

“As a background, Wikimedia Ukraine has been organising Wiki Loves Monuments since 2012, with a goal to collect photos of all cultural heritage monuments of Ukraine on Wikimedia Commons, notably for use on Wikimedia projects. One of the main problems we encountered early on was that our official state lists are biased, especially regarding communist heritage, which was the main reason to have monuments protected or not back when Ukraine was under Soviet rule. And it really matters, as participants are more likely to upload pictures of monuments they associate themselves with.

So we started to organise special categories for less represented monuments of national minorities in Ukraine as early as 2013 — Armenian, Greek, later Crimean Tatar, Jewish, German, Polish, and the most recent one, Bulgarian. It does require more work from us, as we also need to find information about monuments that are not listed officially, or read additional sources to add this or that object to the special categories lists. But it helps us to expand the database of cultural heritage, make it less biased. And it helps us to increase awareness, and motivate veteran participants to continue taking part in the contest, and attract new participants, either interested in the multicultural past of Ukraine, or being a part of those national minorities themselves.

The Polish Heritage in Ukraine campaign was conducted for the second time, and we are very pleased with the results — almost 4600 pictures uploaded by more than 100 authors, depicting 373 monuments, and out of them — 41 are not officially protected by the state, so they are even more endangered, as they can be not only destroyed or damaged by russian drones or rockets, but they can be demolished or repurposed with no oversight by the cultural heritage protection authorities.

The main purpose of a separate special category continues to be to draw attention to these monuments and their condition, and to document them for Wikipedia. And we are very grateful for the support of Wikimedia Polska, that made this project possible, and also to our volunteers and participants for their continued active involvement in the project”.

The Polish Heritage in Ukraine campaign was happening alongside the main contest period for Ukraine in October 2025. Volunteers updated the lists for the special category, so as of now it is containing 1358 monuments (488 out of them with no official protective status). During the campaign itself 102 participants submitted almost 4600 photos, picturing 373 monuments (41 out of them are not registered as monuments) from 15 regions of Ukraine. 24 monuments were pictured for the first time.

A Wiki Loves Monuments Ukraine barnstar

Due to a considerable number of submitted works, there was a pre-selection round. 16 volunteers from Poland took part in reviewing the images. Some Polish volunteers shared their reflections on the photos, their motivation to help, and the process.

Piotr “PMG” Gackowski, Polish volunteer helping with preselection, an editor with almost 9 mln edits on Wikimedia Commons, a Polish Wikipedia administrator, reflected:

“I participated in the preselection of photos for many reasons. One of them is patriotism. In this way, I can support the memory of Poland and the Polish people. The second point is the curiosity typical of every Wikipedian: I took part in the Polish WikiLovesMonuments and wanted to see what photographs from other countries look like. For me, the difference was that the photographs from Ukraine that I was rating much more frequently showed objects in rural areas. In Poland, large cities dominate, so in my opinion it was a significant difference. At the same time, it is important to me that I can help Wikipedians from Ukraine in their work. I am aware that every monument they commemorate by taking photographs could be destroyed”.

Teukros, a Polish pre-selection volunteer, and a Polish Wikipedia administrator, commented: 

“I have participated in the photo preselection process for the Wiki Loves Monuments campaign (Polskie Dziedzictwo w Ukrainie) twice now, and I have genuinely enjoyed doing so. To be honest, I did not have any particularly special reasons for joining this initiative – the simple fact that the Wikimedia community in Ukraine had asked for assistance was reason enough for me.

My experience of participating has been a mixture of sadness and joy. Sadness, because it is plainly visible that Polish heritage sites in Ukraine are often damaged, neglected, and that there is little indication that this situation will improve in the near future. Joy, because the very fact that the Ukrainian community has taken on such a challenge allows us to believe that at least the memory of the Polish presence in these lands will endure.

If I were to say what inspired the greatest sympathy in me during this project, it would, paradoxically, be the photographs that I had to reject. Crooked, overexposed, blurry, taken by amateurs without any special preparation – they were perhaps the strongest testimony that, among completely ordinary people, the memory of the shared history of Poland and Ukraine is still very much alive”.

The pre-selection volunteers reviewed 4571 images (the organising team removed images submitted by the participants with conflict of interest, like organisers and jury members), divided in such a way, that each image was viewed by 3 volunteers.

Archiwald, a Polish pre-selection volunteer, and a Polish Wikipedia administrator, commented:

“In February of this year, I received an offer through WMPL to participate in the preliminary selection process as a person assisting with the initial evaluation of photos. I was happy to join the effort, especially since I already had some experience with similar initiatives at the time. As an editor who focuses, among other things, on historical matters, I realize just how useful the files I’ve been reviewing will be. I’m not just referring to the winning photos here; even those that ultimately didn’t receive any awards add significant value to the Commons resources.

It’s a very pleasant feeling to look through the contest results and notice instances where the judges rated a photo just as highly as I had earlier. I felt that way, for example, when I noticed that the photographs of the palace in Pryozerne by Oleksandr Malyon had been recognized. Photographs like these have immense historical value, which usually becomes apparent only after many years; therefore, the author’s decision to make them available under free licenses deserves recognition”.

Adrian Tync (Gower), another Polish pre-selection volunteer, active on Wikidata, Wikimedia Commons, and Polish Wikipedia, shared: 

“I got involved in the photo pre-selection process because I’d taken part in the ‘Wiki Loves Monuments’ competition a few times myself as a photographer, and I was curious to see what it was like from the other side. I enjoy browsing and evaluating other people’s photos on Commons, for example in the Quality Images nominees section, so this was the perfect task for me. I’m interested in Polish historical monuments and Polish cultural heritage, and thanks to the pre-selection process, I got to see many of them”.

Out of this round 815 photos proceeded to the next round. The organisers reviewed the images more closely, and removed the ones that were not up to the technical standards (like lower resolution), so 711 images proceeded to the main jury, which included Polish Wikimedians and partners of Wikimedia Polska:

  • Damian Kujawa — Wikimedian, volunteer, activist;
  • Julia Szablowska — photo editor, photographer, curator;
  • Magdalena Lachowicz — Assistant Professor at the Department of Eastern Studies at Adam Mickiewicz University in Poznań, Poland;

Each work was viewed by all three jury members. 120 pictures made it to round two, where each jury member was asked to evaluate each work from 1 (minimum) to 10 (maximum) points. The guidance when evaluating pictures was:

  • from 0 up to 3 for technical quality (sharpness, use of light, perspective etc.);
  • from 0 to 3 for usefulness of the image for Wikipedia;
  • from 0 to 3 for originality.
  • 1 additional point for something special in the picture.

The results are presented below, and they are grouped thematically, to showcase the breadth and depth of Polish Heritage in Ukraine, so the awarded works are from different regions of Ukraine, and are grouped by different objects depicted (churches, castles etc). No separate award for active participation — people awarded are among active contributors.

Best photos – Churches (pol. Najlepsze fotografie – Kościoły)

Holy Trinity Church (2022). Velykyi Ostrozhok, Vinnytsia Oblast

The author uploaded the first ever pictures not only of the church, but even from the village itself. And his pictures are now illustrating the article about the village on Wikidata and local Wikipedias (Wielki Ostróżek in Polish Wikipedia, for example). The version in Ukrainian did not even contain the mention of the church, as the building is not a listed monument officially.

Best photos – Palaces, Estates (pol. Najlepsze fotografie – Pałace, Majątki)

Potocki Palace (2025). Tulchun, Vinnytsia Oblast
Rej manor (2025). Pryozerne, Ivano-Frankivsk

Best photos – Other Buildings (pol. Najlepsze fotografie – Inne budynki)

Gymnasium Alexandrinum (2020). Mariupol, Donetsk Oblast

The building was built by a Polish architect Mikołaj Tołwiński, there is no article about the school itself on Polish Wikipedia yet. Due to the Russian occupation of Mariupol, getting new free pictures (or even up to date information about the state of the building) is not going to be a trivial task.

Former house of the Branicki estate manager (2022). Rozkishna, Kyiv Oblast

This is also not a listed building, which makes its status to be more endangered — it is now privately owned, and there were news about it being on sale.

Best photos – Chapels (pol. Najlepsze fotografie – Kaplice)  

Boim Chapel (2014). Lviv

Best photos – Castles (pol. Najlepsze fotografie – Zamki)

Chervonohorod Castle (2020). Ternopil Oblast

Best photos – Residential buildings (pol.Najlepsze fotografie – Budynki mieszkalne)

Alad’yin family house (2025 рік). Kharkiv

During the First World War this house hosted a Polish bookshop, and thus was a cultural center for the Polish community in Kharkiv.

Best photos – Towers (pol. Najlepsze fotografie – Wieże)

Tower on Ford (2025). Kamianets-Podilskyi, Khmelnytskyi Oblast

Best photos – Cemeteries (pol. Najlepsze fotografie – Cmentarze)

A fragment of the wall with names of Polish officers and citizens murdered by the Soviet NKVD in 1940 (2025).Bykivnia graves. Kyiv

Detailed description of each photo in Ukrainian here.

Best video (pol. Najlepsze video)


Saint day in Chornyi Ostriv, Khmelnytskyi Oblast, Ukraine © Philip Gavrilyk (Філіп Гаврилюк), CC BY-SA 4.0
Music “Shadowlands 2 – Bridge” © Kevin MacLeod (incompetech.com), CC BY 4.0

The best video work was decided by a separate jury, consisting of:

  • Roman Barabakh — photographer, traveler, founder of a media project Ukrainian Travels;
  • Oleksandr Havryk — cameraman, editing director, Ukrainian Wikipedian;
  • Maksym Uvaiev — editing director, film critic.

The results of the joint project and winners were celebrated at the Wiki Loves Monuments Ukraine hybrid awards ceremony on May 30, 2026.

At the 2025 Wiki Loves Monuments Ukraine Awards Ceremony 
Polish heritage in Ukraine Special category statistics
Winners present offline
Olena Suhak, one of the winners, commenting on her works
Serhii Plakhotniuk, one of the winners, commenting on his photo
Volodymyr Tarasov, one of the winners, commenting on his contributions
Oleksandr Malyon (on screen), one of the winners 

The winners absent at the event will receive their prizes by post.

Iryna Boiko, communications manager of Wikimedia Ukraine, commented:

“This is the second time we are organising the “Polish Heritage” special category in the Ukrainian edition of the Wiki Lobes Monuments contest. The request to organise a separate category for Polish sites has been repeatedly expressed by the participants themselves, because many such sites are now falling into disrepair and simply collapsing in the absence of an active community to care for them.

The main focus of the special category was churches, Polish cemeteries, castles and fortresses, but the jury also paid attention to residential and other buildings. Of course, the parameters of inclusion in the contest lists are quite wide, because the very idea of ​​the “Wiki Loves Monuments” competition is to collect photos to illustrate Wikipedia articles, and in order for the article to illustrate the life of a community or a certain period, the parameters of inclusion should be quite wide. So, the special category lists contain buildings created by Polish architects for Polish activists, or buildings where Poles lived, or buildings that were important for the Polish community of a particular settlement — like a bookstore in Kharkiv, that became a center of a local Polish community life.

My personal favorite photo, a very symbolic embodiment of what we are trying to achieve through this joint project with Wikimedia Polska, was the work of Valentyn Mahovkin, our long-time participant, which depicts the process of restoring an epitaph on a tombstone in a Polish cemetery in the village of Chornyi Ostriv in Khmelnytskyi Oblast. And, by the way, this cemetery is not officially protected, and only its gate has an official status as a cultural monument…”

Restoration of the epitaph (2024). Polish cemetery. Chornyi Ostriv, Khmelnytskyi Oblast

WikiSuarana: Wikimedia Bandung Community’s Effort to Enrich Content on Sundanese Wikiquote

(WikiSuarana) WikiLatih Wikiquote bahasa Sunda - 18 April 2026 - Komunitas Wikimedia Bandung
Wikiquote training in Bandung (Hasnanf, CC BY-SA 4.0 via Wikimedia Commons)

When people think about Wikimedia projects, most only know the world’s largest online encyclopedia, Wikipedia. Many do not know that Wikipedia has dozens of sister projects, including Wikimedia Commons, Wiktionary, Wikibooks, Wikisource, and Wikiquote. One of the lesser-known projects is Wikiquote. It is a collaborative project that collects and preserves notable quotations from famous people, films or series, fiction and non-fiction books, proverbs, and well-known sayings.

In Indonesia, Wikiquote is currently available in three languages, Sundanese, Banjar, and Gorontalo. In Wikimedia Bandung Community, we believe that community collaboration is essential to enriching Wikiquote’s content and introducing the project to more people. With this goal in mind, we launched WikiSuarana, a community initiative to enrich the Sundanese Wikiquote with quotations related to memorable events and popular trends from 2025.

Project outcomes

WikiSuarana project was carried out by four members of Wikimedia Bandung community: Hasnanf, Raflinoer32, Zulaihamaryam, and Sonofbrahma from March to May 2026. Together, we created 224 Wikiquote articles featuring quotations related to events that took place throughout 2025. Each team member contributed 56 articles to Sundanese Wikiquote.

In addition to our work on Wikiquote, we also contributed to Sundanese Wikipedia by creating 56 articles about notable events from 2025. We know that many Sundanese speakers still use Sundanese Wikipedia as a source of information. By enriching Wikipedia with these articles and linking them to the related Wikiquote pages, we make it easier for readers to discover and explore our collection of quotations on Sundanese Wikiquote.

Community outreach through training and meets up

Community outreach for Wikisuarana was carried out during the month of April 2026, consisting of a series of training and meet-up activities. As a warm-up activity, in the first week of April, we conducted a community meet-up to edit on Sundanese Wikipedia together focusing on creating new articles regarding remarkable events that happened throughout the year 2025 and the notable figures related to it. This event was attended by 13 participants, both online and offline, which created 14 new articles on Sundanese Wikipedia. By starting WikiSuarana with this thematic edit activity, it was expected to provide a thematic context regarding the purpose of this project, which is to document remarkable events and notable figures in Sundanese Wiki projects. Some of the articles that were the results of this event, for example Tambang Grasberg, Satelit Nusantara Lima, and Sri Rejeki Isman.

On April 18, the Asia-Africa Conference is annually commemorated in Bandung, since the city was the first host of the conference back then in 1955. On that day this year, we conducted a Wikiquote training activity partnering with an independent library in Bandung, which carried out the theme of anti-colonialism spirit. In this activity, we focused on training new contributors to edit and create new quotation articles on Sundanese Wikiquote regarding anti-colonial figures, anti-colonial literatures, and those related to the Asia-Africa Conference. The training was attended by 13 participants, which created 15 new articles. Hopefully, other than attracting new contributors, this event could be the beginning of a consistent effort to amplify and document the voice of anti-colonial figures through Sundanese Wikiquote. Some of the articles that were the results of this event, for example Behind The Scenes – Story of The Bandung Conference Committee, Teh dan Pengkhianat, and Leila Khaled.

The series of WikiSuarana was concluded with another community meet-up as a continuation of the prior training, which was focused to create and edit quotation articles on Sundanese Wikiquote regarding the theme of anti-colonialism and remarkable events that happened during the year 2025. This event was attended by 11 participants, both online and offline, which was also attended by some of the participants of the previous training event, and produced 14 new articles. Some of the articles that were the results of this event, for example Jawaharlal Nehru, Mohammad Yamin, and Ali Sastroamidjojo.

Lesson learned

Through WikiSuarana project, we all learned two valuable lessons:

The first is the importance of partnerships. As a local community, collaborating with organizations that share our mission is essential to promoting free knowledge, especially about the Sundanese language and culture. These partnerships help introduce Wikimedia projects and our community to a wider audience. During this project, we collaborated with an independent library in Bandung. In the future, we hope to work with more partners to organize Wikimedia activities such as workshops, research projects, and community meetups.

The second lesson is about promotion and outreach. In today’s digital world, many people get information through social media platforms such as Instagram and TikTok. Throughout the project, we created promotional content, from the project launch to activity announcements. However, we learned that relying on just one or two social media platforms is not enough. Recently, Threads has become increasingly popular, and one of our team members found that project posters shared there reached a wider audience and received positive engagement. Based on this experience, Wikimedia Bandung plans to use Threads alongside our other social media channels to promote future Wikimedia activities.

What’s next?

Documenting and amplifying the voices of the people who are part of history is a way to preserve our collective memory. By narrating them in their own voices, hopefully we can make sure that the history being told is honest-to-goodness. Through contributing it to Wiki projects, especially Sundanese Wikiquote, we also hope to be able to preserve them in our mother tongue.

Our plan forward is to keep documenting other voices that are still unheard while also promoting the sister projects of Wikipedia, which already has the Sundanese version, such as Wikiquote. We also plan to reach outward to other parts of West Java to attract many other contributors so that our community can keep growing while also adding much other knowledge to the Wiki projects itself.

Hatur nuhun!

Recap: Wiki Loves Pride Lagos Physical Event and Lessons from the Past Six Months

Wiki for Human Rights, Nigeria - Wiki Loves Pride
Wiki Loves Pride Flyer – Lagos

I started working as a volunteer community manager for Wiki for Human Rights, Nigeria, in November 2025 with a focus on organizing trainings and retaining LGBTIQ+ Nigerian editors as active contributors to Wikimedia projects. Although I created my Wikimedia account on 1 March 2024, I did not understand how editing worked and never made any contributions after creating it. 

The same year, I attended a Wikimedia session on Queerpedia. As a writer who is passionate about volunteering, particularly in the open knowledge movement, I still left without knowing how to contribute. The session ended with participants creating accounts but without a practical understanding of how Wikimedia actually worked. This is something I have observed among newcomers: navigating Wikipedia, the most well-known Wikimedia project, can be a daunting experience.

Wiki Loves Pride - Lagos 2026
After-session group picture

That changed when I attended Wiki Loves Pride 2025, held on 29 June. It was my first in-person Wikimedia event, and with my laptop beside me, everything about contributing to Wikipedia suddenly became much clearer. Through editing Wikipedia, I also discovered the wider Wikimedia ecosystem and its sister projects. Even after the training, I still encountered a few challenges, particularly with adding awards, information tables, and infoboxes. However, through conversations on WhatsApp with the Wikimedia Nigeria Project Officer, Ayokanmi Oyeyemi (user: Kaizenify), who facilitated the training, as well as guidance from Wikipedia help pages and Wikimedia Commons documentation, I was able to overcome those challenges.

Since assuming the role of Community Manager and Project Officer for Wiki for Human Rights, Nigeria, I organised monthly virtual and physical training sessions aimed at improving editor retention. Organising both the physical and virtual Wiki Loves Pride campaign this June felt like a full-circle moment; déjà vu. It also became an opportunity to reflect on the learning experiences from reviewing participants’ contributions and identifying areas where new editors commonly struggled. 

Wiki Loves Pride - Lagos 2026
Tony Obinna facilitating a session

Since becoming an active Wikimedia contributor in July 2025, I have made more than 3,000 edits across Wikimedia projects, with a primary focus on LGBTIQ+ and Nigerian topics. Beyond editing, I have also taken on leadership positions, including serving as a committee member for the Nigerian National Funding program and as a core organizing member for Queering Wiki, scheduled to take place later this year in Canada.

Alongside the online Wiki Loves Pride campaign, we partnered with the Centre for Population Health Initiatives (CPHI),  a health organization that serves both the general population and minority communities, to host Wiki Loves Pride. The program combined a Pride celebration with Wikimedia training for both experienced and new editors. It also marked the first time I independently facilitated an entire Wikimedia training session from the beginning to the end: account creation, making edits, and introducing participants to the broader Wikimedia ecosystem.

Wiki Loves Pride - Lagos 2026
Participants engaged throughout the session

As someone who has always dreaded public speaking because of a minor speech impediment, becoming part of the Wikimedia movement as a community leader has helped me grow tremendously. Standing in front of more than 15 participants and leading a session on documenting queer knowledge, I did not freeze or lose confidence. Community advocacy for LGBTIQ+ people in Nigeria has always required me to speak publicly from time to time, but Wikimedia has made it a consistent part of my work through monthly virtual and physical trainings. It has taught me that confidence in public speaking is often built through practice, and that many of the fears we carry can gradually be overcome through repeated experience.

Grateful and Growing: Reflecting on My Wiki Afrodemics Fellowship, Pilot Cohort

By: Ogundele1

When the Wiki Afrodemics Mentorship Programme kicked off in March 2026, I knew I was stepping into something special, what started as a curiosity to learn more about Wikimedia projects has turned into a transformative journey that has completely reshaped how I contribute to free knowledge.

The Wiki Afrodemics Mentorship Programme assembled 20 passionate participants from underrepresented African countries, super proud to be one of the selected participants from the pool of over 300 applicants. This project empower fellows through structured training sessions, hands-on editing, and collaborative activities across Wikipedia, Wikidata, and other Wikipedia sister projects. Being part of this diverse community of learners was inspiring, we came from different countries, spoke different languages, but shared a common mission.

Wiki afrodemics and mentorships programme mentees

Enhancing My Wikidata Skills

A personal highlight of this programme was the significant improvement in my Wikidata editing skills, made possible through the exceptional facilitation of my favourite mentor, David Partey. While I was already familiar with Wikidata , Ialways believe there’s more to improve on; structured data requires a distinct mindset and technical approach.

David’s training sessions were invaluable. He broke down complex concepts like creating new items, adding statements with reliable references, and querying data from the Wikidata platform. His patient and structured approach demystified Wikidata, turning it from a daunting database into an intuitive and powerful tool for enhancing the visibility of African academics.

Under his guidance, I learned not just how to edit Wikidata, but why it matters. I now understand how to;

  • Add meaningful statements with reliable references
  • Connect Wikidata to Wikipedia articles and Wikimedia Commons files
  • Use Wikidata to make African academics more visible online.

This newfound proficiency has made me a more confident and well-rounded editor, capable of contributing meaningfully across multiple Wikimedia projects. Today, I can confidently say that my Wikidata editing skills have improved tremendously, and I owe so much of that growth to David’s exceptional facilitation.

The sessions facilitated by other mentors  were equally impactful. Each mentor brought unique expertise and perspectives, and I soaked up every bit of knowledge they shared. The collaborative atmosphere, the peer feedback, and the sense of community made learning feel less like a classroom and more like a family gathering.

Up next!

As this maiden Cohort wraps up, I am filled with so much gratitude for the mentors who invested their time and expertise in us, for the facilitator who believed in this vision, and for my fellow fellows who made this journey so memorable.

The Wiki Afrodemics Mentorship Programme has shown me that mentorship is not just about receiving, it is about growing, connecting, and ultimately giving back. I am leaving this programme as a Wikipedian, a more knowledgeable Wikidata contributor, and a passionate advocate for free knowledge in Africa.

I cannot wait to apply everything I have learned and to contribute to future cohorts, this time, not as a fellow, but as someone who can support and inspire others just as I was supported and inspired.

Thank you, Wiki Afrodemics, for this life-changing opportunity. This is just the beginning of my journey.

Wikimania 2026: Freedom, equity and reliability

How can the Wikimedia community defend freedom, equity, and reliability on the internet? Are we in a global information crisis?

Several sessions on Day 2 of Wikimania Paris explored these questions from different angles. The morning kicked off with breakout sessions on how global trends are impacting government regulation – with attendees joining Wikimedia Foundation board members in small group discussions. Protecting free knowledge was a core theme throughout the day.

Misinformation on climate change

During the 2025 Iberian peninsula blackout, misinformation was spread that the main cause was renewable energy. In “Climate Conversations: Interdisciplinary Approaches to Knowledge Sharing”, moderator Tatjana Baleta said that rapid response editing on Wikipedia was one way that the scientific community combats misinformation during extreme weather events and other major incidents.

Many readers have also shifted from reading about climate change to reading more about other issues such as the cost of living crisis. Dr. Femke Nijsse (User:Femke) discussed the importance of meeting readers where they are – and explaining the science related to these issues.

While AI-generated content may have the sheen of reliability, it often turns out that the sources they cite are hallucinations or that they do not actually verify the claims made by the LLM. Editors on Wikipedia are now starting to use tools such as AI Source Verification (which itself uses LLMs) to predict the verifiability of claims made in a article.

AI crawlers: Encroaching on creativity?

“Collateral Damage? Human Creativity and Interaction in the AI Crawling Era” explored how organizations in the free knowledge ecosystem are responding to the massive increase in AI scrapers.

There was a consensus among panelists that attribution is a critical concern for authors. Creative Commons CEO Anna Turnadóttir and Monica Westin of Cambridge University Press discussed the need to educate authors about the benefits of open access models – while also acknowledging the need to experiment.

Mark Graham discussed how the Internet Archive is reaching out to news organizations to discuss alternatives to blocking the Wayback Machine, such as rate limiting and allowing access only for certain uses.

Striking a balance

Throughout the day, speakers debated the complexities of regulation – and how the rush to “do something” can backfire. Turndóttir said, “What used to be the internet handshake online is now the middle finger. And that sort of environment forces lawmakers to reach for blunt tools like regulation. The community needs to establish norms, but sometimes regulation does cause real harm.” 

The keynote session on “Protecting Free Knowledge – The New Battlegrounds of Digital Freedom”, Nnenna Nwakanma (from the internet) discussed the discourse around regulation – and how European approaches may not be applicable in Africa. Panelists during the session, including Axelle Lemaire, architect of the 2016 loi numerique, emphasized the need for balance in protecting openness and freedom online, while also protecting the privacy of individuals.

The conference is making waves in France with media coverage in 25 outlets – including La Croix‘s print edition, a Radio France podcast, and more.

Indic Wikimedia Hackathon Hyderabad 2026

Group photo of the Indic Wikimedia Hackathon Hyderabad 2026, Image by Nivas

Program Purpose

Fifty-six contributors gathered at IIIT Hyderabad for three days to improve Wikimedia’s technical ecosystem. Unlike traditional hackathons that focus primarily on rapid prototyping, the Indic Wikimedia Hackathon 2026 introduced dedicated refinement sessions that encouraged participants to improve code quality, documentation, usability, and long-term maintainability.

The Indic Wikimedia Hackathon Hyderabad 2026 was organized by Indic MediaWiki Developers User Group (aka Indic-TechCom). The hackathon took place in Hyderabad from 26 – 28 June 2026 (with 25 June as Day 0), in collaboration with the Open Knowledge Initiatives team and OSDG club at International Institute of Information Technology, Hyderabad.

Wikimedia hackathons are spaces for developers, designers, content editors, and other community stakeholders to collaborate on building technical solutions that help improve tools, workflows, and overall user experience across Wikimedia projects.

This hackathon is designed for:

  • Technical contributors active in the Wikimedia technical ecosystem, which includes developers, maintainers (admins/interface admins), translators, designers, researchers, documentation writers, etc.
  • Content contributors having an in-depth understanding of technical issues in their Wikimedia projects, like Wikipedia, Wikisource, Wiktionary, etc.
  • Contributors to any other open-source community or those who have participated in Wikimedia events in the past, and would like to get started with contributing to Wikimedia technical spaces.

Participants worked on a curated set of technical tasks prepared by mentors and organizers. They were also encouraged to propose their own project ideas, provided they included a clear problem statement, implementation approach, and were reviewed by mentors before the event. 

Building on the experience and learnings from previous hackathons, this event was more efficient, inclusive, and collaborative. 

The event aimed to involve more developers who have experience with the Wikimedia ecosystem and had prior experience already contributing to tools, extensions, gadgets, or other technical projects. Editors were paired up with developers to provide domain knowledge, helping them better understand editing workflows, user needs, and the intended behaviour of the applications/ extensions/ gadgets being developed. 

Unlike other hackathons where rapid development is the primary focus, this event was not completely hacking but also incorporated dedicated  refinement sessions.These sessions encouraged participants to improve the quality of the works by refining  design, UI, data privacy, code optimization, documentation and overall maintainability. Additionally, the program also included brainstorming sessions, group discussions, workshops  to help participants  understand a broader perspective of this ecosystem beyond their individual projects.

Scope and Timeline

The hackathon was conducted as a three day in-person event  at the International Institute of Information Technology, Hyderabad (IIIT-H), a long-standing partner that provides space for technical and community events. The venue supported collaborative work through dedicated hacking spaces, mentor interactions, and discussion areas. 

The scope of the event was flexible enough to encourage participants to work on a curated set of Wikimedia-related technical projects and tasks suitable for a hackathon which prepared by mentors or a custom project proposed by their own and reviewed by experienced developers and organizers, along with Team Challenges from Wikimania Hackathon 2026.

The program was structured into distinct phases. The first half of the event (approximately one and a half days) focused entirely on development where the first part of the first day was catered to some introductions and welcome notes, ground rules, ice breaker activities, followed by continuous hacking. On the second day, the event had a social activity and the morning session was mostly catered to workshops, brainstorming sessions, and getting to know about OKI work. The second half started with refinement phases. The third day focused completely on refinement, wrap-up and showcase. 

Attendance

Total Attendees: 56

  • Organisers:   11
  • Mentors:   15
  • Participants:  25
  • Editors:  5

The majority of participants were developers with prior experience in the Wikimedia technical ecosystem. A smaller group consisted of experienced Wikimedia editors with technical knowledge, who collaborated with developers by providing domain expertise and user perspectives during the hackathon. 

Activities Conducted

Prior to the hackathon, an orientation call was conducted for participants to give an overview of the program, the Wikimedia technical ecosystem, team formation details, and some logistical and operational arrangements. The session also introduced participants to Wikimedia, its technical ecosystem, guided them through basic account setup, and explained the overall hackathon format. 

Following the orientation,  participants were encouraged to have a call with their specific teams and mentors. These discussions helped participants better understand their assigned projects, including the scope, objectives, expected outcomes, and technical requirements. Mentors introduced project-specific workflows, outlining the scope, objectives, and tasks for each participant to help them engage effectively during the hackathon.

During the event, participants worked on pre-curated tasks across multiple Wikimedia-related projects and collaborated closely with mentors to understand issue tracking, patch submission, and debugging workflows. Mentors supported participants across different projects, helping them navigate both technical challenges and Wikimedia-specific contribution processes.

To complement the technical program, the hackathon also included community-building activities. An icebreaker session at the beginning of the event helped participants interact and build connections. On the second day, interested participants joined a social walk around Hyderabad, providing an informal opportunity for networking. A dedicated women’s dinner was also organized to foster stronger connections among women participants, encourage inclusion, and support long-term retention within the Wikimedia technical community. 

Towards the conclusion of the event, a project showcase and presentations session was conducted, during which participants demonstrated their work and shared learnings with fellow participants, mentors, and organizers.  

Outputs and Outcomes

The hackathon enabled participants to work on pre-identified tasks across multiple Wikimedia-related repositories, resulting in code contributions, feature enhancements, bug fixes, and documentation improvements. Throughout the event, participants gained practical experience with Wikimedia development workflows, including issue tracking, patch submission, code review, and collaborative problem-solving, with guidance from experienced mentors. Several participants continued engaging with their assigned projects after the event, indicating effective onboarding into Wikimedia technical workflows.

During the hackathon, participants submitted a total of 56 Phabricator tickets across multiple Wikimedia-related projects. The distribution of contributions is summarised below:

  • Clip2Commons:  9
  • Deployr: 7
  • Language Selector Rewrite: 1
  • Lingua Libre: 5
  • Montage:  4
  • NPOV Drift Detector: 1
  • Observability Tool: 7
  • Onboarding of New Wikipedia Editors  : 2
  • OpenSpeaks Subtitler: 5 
  • OpenSpeaks Tome: 3
  • Pywikibot  : 3
  • Scribe: 6
  • Translate Tagger: 8
  • ULS Extension :8
  • Wanda Extension: 4
  • WikiEval Tool : 3
  • Wikievol Tool  : 3
  • WikiLinkua  : 2
  • Wikimedia Commons Android: 4 
  • Wikisource Reader App: 2

Total: 87 Repo/Phabricator tickets linked

As part of the Indic Wikimedia Hackathon Hyderabad 2026, several teams worked on projects that directly align with the official Team Challenges announced for Wikimania 2026. 

  • OpenSpeaks (Subtitler/Tome/Bento) – Boost multimedia experience, Connect multilingual knowledge
  • Lingua Libre – Boost multimedia experience, Connect multilingual knowledge
  • WikiLinkua – Gamify knowledge, Stream data with Wikidata
  • Wiki Translate Tagger – Connect multilingual knowledge
  • Wanda / WandaScore / WandaScribe – Welcoming newcomers, Fix the sources / Update the obsolete, The editor of the future
  • Scribe – Stream data with Wikidata, Connect multilingual knowledge
  • WikiEvolution – Explore knowledge
  • WikiNPOV Drift Detector – Deciphering biases
  • Language selector rewrite – Connect multilingual knowledge

These contributions included code changes, improvements, and related updates submitted under mentor guidance.

What went well:

Program design:

  • The overall event design enabled participants to collaborate with peers from diverse backgrounds and work effectively on technical projects.
  • The combination of structured onboarding, continuous hacking, and dedicated refinement sessions supported steady progress throughout the event.
  • Refinement sessions encouraged participants to improve code quality, documentation, user interface, and maintainability rather than focusing solely on completing tasks.

Mentorship and technical contributions:

  • Clear project introductions and continuous mentor support helped participants engage confidently with their assigned tasks.
  • Most of the identified hackathon tasks were actively worked on during the event.
  • Several participants continued contributing to their assigned projects after the hackathon, demonstrating successful onboarding into Wikimedia technical workflows.

Collaboration:

  • Pairing editors with developers proved valuable, as editors helped developers better understand user workflows, expected tool behaviour, and usability considerations.
  • Workshops and discussion sessions complemented the hacking sessions by providing participants with a broader understanding of the Wikimedia technical ecosystem.

Diversity and inclusion:

  • The hackathon achieved approximately <>% women participation, the highest among events organized by the User Group to date.
  • The women’s dinner helped foster stronger connections among women participants and contributed to a more welcoming environment.

What can be improved/Learnings?

Program schedule

  • Participants requested additional time for the project showcase and presentations.
  • Fifteen-minute breaks were considered too short and could be extended in future editions.
  • Starting sessions at 9:00 AM posed challenges for some participants because of commuting time.

Project selection

  • Hackathon tasks require more review before the event to reduce duplication of effort.
  • Greater emphasis should be placed on improving existing Wikimedia tools rather than developing new ones where similar solutions already exist.

Mentorship

  • Participants experienced delays when waiting for scheduled online mentor support.
  • Increasing mentor availability or ensuring more mentors are physically present during the event could improve the overall experience.

What’s next:

Based on the outcomes and observations from the Indic Wikimedia Hackathon Hyderabad 2026, the following recommendations are proposed for future editions of the event:

  • Prioritise improving existing Wikimedia tools and applications over developing new ones, where appropriate.
  • Continue incorporating dedicated refinement sessions to improve the quality and sustainability of contributions.
  • Allocate more time for project showcases, presentations, and participant discussions
  • Review the event schedule by extending break durations and considering a later start time where feasible.
  • Continue initiatives that promote diversity and inclusion, including activities that support participation and retention of women contributors.
  • Strengthen post-event follow-up and mentorship to encourage continued contributions beyond the hackathon.
  • Explore organizing additional hackathons, workshops, and technical events in college campuses to improve accessibility and encourage new technical contributors.

From a Wild Idea to Reality: Building WISE at Wikimedia Hackathon 2026

Wikimedia_Hackathon_2026_group_photo_03
Wikimedia_Hackathon_2026_group_photo_03

When people think about hackathons, they often imagine coding sessions, project demos, and late nights spent debugging. For me, Wikimedia Hackathon 2026 in Milan was something much more meaningful: a reminder that some of the most impactful ideas begin as simple conversations between passionate people.

Like many Wikimedia Commons contributors, I have often found it difficult to discover media through traditional search. Commons hosts millions of images, videos, and audio files, but today’s search relies mostly on filenames, categories, descriptions, and structured data not on what is actually visible or audible in the file itself. Hundreds of campaigns, contests and content initiatives add new media every year, which only makes the problem bigger. I kept wondering, half as a joke and half seriously: what if we could search Commons based on what an image or video actually shows, rather than what someone happened to type into its metadata?

It felt like a wild idea at the time, so I shared it with the Wikimedia technical community mostly to see if anyone else found it interesting.

Eugene, David and Gopa at WMHACK 2026
Eugene, David and Gopa at WMHACK 2026

A researcher and developer named David, who had recently joined the Wikimedia community, came across the discussion. David had been working on a technology called WISE.. at the University of Oxford research focused on semantic understanding of visual and multimedia content, enabling people to search images, videos, and audio using natural language. What started as a random online exchange quickly turned into a collaboration. David hadn’t originally planned to attend the hackathon in Milan, but as we kept talking, we both got excited about bringing this kind of search to the Wikimedia ecosystem, and decided to meet in person and try to build it.

Looking back, that decision changed everything.

For two intense days at the hackathon, we worked side by side to integrate and demonstrate WISE for Wikimedia Commons — architecture, datasets, search quality, the usual string of small technical fires. One moment that stuck with me: the first time we typed “horse in an airplane” into the prototype, half-expecting nothing, and watched it actually return the right video. That was the point it stopped feeling like a demo and started feeling like a real tool.

What inspired me as much as the technology was David’s approach to problem-solving. In an era where most developers reach for an AI assistant the moment something breaks, David would just as often dive into technical documentation, research papers, and manuals first. Watching him methodically work through a problem rather than skip to an answer was a good reminder that strong engineering fundamentals haven’t gone anywhere.

What we built

The result of those two days is WISE, a new experimental search experience for Wikimedia Commons. It currently indexes Media of the Day content around 5,000 videos and is built to understand the actual visual and audio content of a file rather than its metadata.

Project Wise showcasing Semantic Image Search
Project Wise showcasing Semantic Image Search

Semantic search. Search using natural language and find media based on what appears in the image or video itself. A few examples that worked surprisingly well:

  • “man at a train station”
  • “horse in an airplane”
  • “man with a flower”
  • “pirate with a pistol”

Audio search. Search within audio content to find relevant recordings and segments.

Project Wise showcasing Facial Recognition Feature
Project Wise showcasing Facial Recognition Feature

Face search. Upload a photo of a face, and WISE can locate that person across images and videos… for video, it can even surface the timestamps where they appear.

Multilingual search. Queries work in multiple languages, including Hindi and Telugu. This matters more than it might first appear: most of Commons’ existing search tooling is built around English-language metadata, which quietly shuts out a large share of Wikimedia’s global, non-English-speaking contributor and reader base. A search experience that understands a query in Hindi or Telugu as well as it understands one in English is a small but real step toward making Commons more usable for the movement it actually serves.

You can try WISE yourself here: wise.wmcloud.org
Commons Project page: commons.wikimedia.org/wiki/Commons:Wise

What’s next

This is only the beginning. We’re already discussing:

  • Expanding indexing beyond Media of the Day to cover Commons at a much larger scale.
  • Finding visually similar images after an upload.
  • Suggesting categories, filenames, and metadata based on visual similarity.
  • Improving search quality and broadening multilingual support further.

Most of all, this experience reminded me why I love being part of the Wikimedia movement. A passing idea shared on a community forum connected two people from different backgrounds and different parts of the world. An online discussion became an in-person collaboration. A concept became a working prototype. And a hackathon became the place that vision came to life.

Sometimes the most valuable outcome of sharing an idea isn’t the idea itself it’s the people who find it, connect with it and decide to build something together. For me, WISE for Commons is more than a search tool. It’s proof of that.

What Train the Trainer 2026 Taught Me About Building Stronger Wikimedia Communities

When I applied for Train the Trainer (TTT) 2026, I expected to learn how to organize better events, facilitate workshops, and become a more effective trainer.

Over three days in Hyderabad, I certainly learned those skills but I also came away with something far more valuable: a new understanding of how strong Wikimedia communities are built and sustained.

Rather than focusing only on editing or technical skills, the program explored the people behind Wikimedia the contributors, organizers, mentors, and volunteers who make free knowledge possible. Through discussions, hands-on activities, and collaborative exercises, TTT encouraged participants to think beyond individual contributions and toward building welcoming, resilient communities.

Here are some of the lessons that stayed with me long after the program ended.

Communities come before content

One of the biggest surprises on the first day was that very little time was spent talking about editing, and all.

Instead, sessions explored trust, belonging, leadership, contributor motivation, and community participation.

The panel discussion “What Makes Us Stay? Trust and Participation in Communities” highlighted something every Wikimedia community experiences: attracting contributors is only the first step. Helping people feel welcomed, supported, and valued is what encourages them to stay.

Another activity challenged participants to analyze contributor data across language communities. Looking at the gap between the large number of people who consume knowledge online and the relatively small number who actively contribute made me rethink community growth. I realized that successful outreach is not measured only by how many people join an event it is also measured by how many continue contributing afterwards.

The sessions on leadership, trust, and the Universal Code of Conduct reinforced another important message: healthy communities are built intentionally. Trust is earned through respectful collaboration, inclusive spaces, and consistent support for newcomers.

Wikimedia Commons is about preserving knowledge not just photographs

Before attending TTT, I believed I already understood Wikimedia Commons.

I had uploaded photographs, organized Wiki Science competitions, and knew the basics of licensing and file uploads.

The Commons sessions completely changed that perspective. Rather than focusing on uploading more images, the discussions emphasized documenting knowledge in ways that remain useful for future contributors. We explored why metadata, categories, descriptions, geolocation, and licensing all play an essential role in making media discoverable and reusable. One idea particularly stayed with me. Commons does not necessarily need another photograph of a monument that has already been documented hundreds of times. It needs photographs that document the subject well.

That simple idea changed the questions I ask before uploading an image. Instead of asking whether I can upload a photograph, I now ask whether it genuinely helps someone understand a place, object, or tradition better.

The photo walk around the IIIT Hyderabad campus gave participants an opportunity to apply these ideas immediately. Rather than simply taking attractive photographs, we practiced documenting subjects from angles that communicated information clearly and added educational value.

The licensing session also helped demystify Creative Commons licenses through practical examples, making it easier to understand how open licensing enables collaboration across Wikimedia projects.

Good communication keeps communities growing

The final day focused on communication an area that is often overlooked but essential for sustaining volunteer communities.

A session on Visual Storytelling in Practice demonstrated how photographs and personal stories can help communicate knowledge more effectively than facts alone. It reminded me that contributors often remember stories long after they forget presentations.

Another workshop explored communication strategies for different audiences. Working in groups, participants designed outreach plans tailored to specific communities rather than relying on a single approach for everyone. That exercise reinforced a simple but important lesson:

There is no single Wikimedia audience. Students, teachers, heritage enthusiasts, language learners, and professionals all engage with Wikimedia for different reasons. Effective outreach begins by understanding those motivations.

The session on Wikivoyage also introduced me to another Wikimedia project that I had not previously explored in depth. It demonstrated how documenting travel knowledge, local culture, and places contributes to the broader free knowledge ecosystem.

Finally, a session on community communication platforms highlighted how mailing lists, Telegram groups, discussion forums, and social media help communities stay connected long after events have ended.

Building a community is not only about organizing an event. It is about creating ongoing conversations.

Learning by doing

One aspect of TTT that I particularly appreciated was its emphasis on practical learning.

Rather than relying entirely on presentations, participants engaged in group discussions, storytelling exercises, communication planning, case studies, and hands-on Commons activities.

These exercises encouraged us to apply ideas immediately instead of simply listening to them.

The collaborative nature of the program also created opportunities to learn from participants representing different language communities, projects, and experiences across the Wikimedia movement.

That diversity of perspectives became one of the most valuable parts of the training itself.

What changed after TTT?

The impact of Train the Trainer extended well beyond the three days of the program.

It changed how I think about community building.

Instead of focusing primarily on organizing events, I now think more about contributor retention, mentorship, and creating welcoming spaces for newcomers.

It changed how I contribute to Wikimedia Commons.

I now pay much greater attention to documentation quality, metadata, licensing, and preserving local heritage through meaningful photographs.

It also influenced my later work within the Wikimedia movement. Many ideas that I developed while preparing proposals for WikiConference India, as well as my growing interest in documenting local heritage and strengthening community communication, can be traced back to discussions and activities during TTT.

Most importantly, the program encouraged me to think beyond individual edits and toward strengthening the communities that make those edits possible.

Why programs like Train the Trainer matter

Every Wikimedia community faces different challenges, but many of those challenges share common themes: welcoming newcomers, retaining contributors, building trust, documenting knowledge responsibly, and communicating effectively.

Train the Trainer creates a space where volunteers can exchange experiences, learn from one another, and return home with practical ideas that can be adapted to their own communities.

For me, the greatest takeaway was not a single workshop or activity.

It was a shift in perspective.

Wikimedia is sustained not only by articles, photographs, or software, but by people who collaborate, mentor, listen, and continue learning together.

That is what I brought home from Train the Trainer 2026 and it is why I believe programs like TTT continue to play an important role in strengthening the Wikimedia movement.

Wikimedia Foundation Bulletin 2026 Issue 13

Here is a quick overview of highlights from the Wikimedia Foundation since our last issue on July 3. Previous editions of this bulletin are on Meta. Let foundationbulletin@wikimedia.org know if you have any feedback or suggestions for improvement!

Highlights

  • Reflections from around the puzzle globe: Wikimedia Foundation CEO, Bernadette Meehan shares her reflections from around the puzzle globe.
  • Wikimania 2026Wikimania is happening this week! After the event, all streamed sessions will be linked in the program on Eventyay and later uploaded to Commons.
  • Grantmaking: The Global Resource Distribution Committee has published a Grantmaking Strategy draft that sets out a renewed approach to how the Wikimedia Foundation’s Community Fund is distributed across the Wikimedia Movement. The GRDC is requesting feedback from volunteers and affiliates, regardless of whether they are grantees or not.
  • Movement EcosystemA proposal that would update movement affiliate recognition and establish new, connected criteria for eligibility to receive Community Fund grants is now available for community review. You are invited to read the proposal and participate in the discussion until August 7.
  • AI’s impact on Free Knowledge: How AI threatens the social contract of free knowledge and what we can do about it.

Annual Goals Progress on Engage

See also: Growth · Product Safety and Integrity · Tech News · Language and Internationalization · The Wikipedia Library · list of movement events · Wikifunctions & Abstract Wikipedia

  • The Wikipedia LibraryFive new collections were added to the Wikipedia Library and a course developed by leading experts in academic publishing and open knowledge was launched.
  • Women+ contributions in Wikimedia TechA guide based on lived experience on how to address some of the invisible barriers for women+ in more technical Wikimedia spaces and recommendations to become more inclusive. Help further by filling out this survey to better understand technical contributions by women+ across Wikimedia projects until July 20.
  • Structured ExperimentationA reflection on the first year of structured experimentation highlights successful experiments such as Paste Check, Reference Check, and Tone Check, which improved editing outcomes and have been rolled out to more users, as well as experiments that did not lead to product changes.
  • Revise Tone test ended: The A/B test of Revise Tone ended on July 9. It showed that newcomer task completion rates increased by 38.7% compared to the default Copyedit task, with no decrease in edit quality. The feature is now available for everyone on the Arabic, English, French, and Portuguese Wikipedias. The plan is to release Revise Tone to more wikis.
  • Keeping mentor list up to date: Administrators on wikis where Growth features are available can now automatically remove inactive mentors by configuring the settings at Special:CommunityConfiguration/Mentorship to keep the list updated. Mentors are experienced contributors who opt in to help new users on-wiki through the Growth Features.
  • WikifunctionsHow Abstract Wikipedia fits into the Wikimedia Foundation’s annual plan FY26/27.
  • Wikidata: The latest Wikidata Platform newsletter (July edition) shares how to identify and rewrite queries that rely on Blazegraph-specific extensions and affected by the migration off Blazegraph.
  • Discussion Tools: On English Wikipedia, DiscussionTools‘ Usability Improvements has now become default for talk pages. You can opt-out of these changes at any time in user preferences. With this, Discussion Tools are now fully available at all wikis.
  • Tech News: The latest highlights from Tech News week 28 and 29 include the new Parsoid parser continues to be deployed to additional wikis, making it easier to introduce new reading and editing features. See also the 72 community submitted tasks that were resolved over the last two weeks. Overall, from April – June 2026 about 337 community tasks were resolved by the Wikimedia Foundation.
  • Language inclusivity at Wikimania 2026New approaches to translation and interpretation will be tested at Wikimania this year.
  • Call for submissions open for Wikimedia Latin America Conference 2026: The Wikimedia community in Latin America has opened the call for session proposals for the Wikimedia Latin America Conference 2026. Community members are invited to submit proposals for the conference program by August 10.

Annual Goals Progress on Protect

See also: Global Advocacy blog · Global Advocacy Newsletter · Policy blog

  • AdvocacyChina again blocks the Wikimedia Foundation as a permanent observer to the World Intellectual Property Organization (WIPO).
  • UK Online Safety Act: Ofcom, the United Kingdom’s Office of Communications, announced that Wikipedia is not designated as a Category 1 service under the Online Safety Act (OSA). This is an important and welcomed outcome as a Category 1 designation could have included privacy and safety risks to our global community of volunteers.
  • UN Open Source Week edit-a-thon: Volunteers created 60 new Wikipedia articles and made nearly 700 updates to improve Wikipedia’s coverage of UN and open source topics at the second UN Open Source Week edit-a-thon co-hosted by Wikimedia Foundation.
  • “Don’t Blink”The latest developments from around the world about protecting the Wikimedia model, its people and its values.

Annual Goals Progress on Reach

See also: Wikimedia Apps · Readers

  • “A Wiki Minute” videosNew videos are added to the series answering some of the most common questions such as “Do you still need Wikipedia when AI can answer anything?” and “Does Wikipedia push a political agenda?”.
  • Wikipedia 25 brand collaboration in Indonesia: On 4 July, the Jakarta-based street wear company Ageless Galaxy launched a Wikipedia 25 collection, the first ever Wikipedia Brand Collaboration in Asia. The apparel collection featured hats, t-shirts, and a jigsaw puzzle cardigan. Wikimedians in Indonesia joined Ageless Galaxy for a launch party.

Other updates

Board and Board committee updates

See Wikimedia Foundation Board noticeboard · Affiliations Committee Newsletter

  • Board Elections Eligibility: There are new proposed eligibility criteria for standing as a candidate in Wikimedia Foundation Board of Trustees election now available for feedback. The requirements are more specific and detailed than in years past, to both inform the community of what the Board needs and to create multiple pathways to the Board for Wikimedians.

Other Movement curated newsletters & news

See also: Diff blog · Goings-on · Planet Wikimedia · Signpost (en) · Kurier (de) · Actualités du Wiktionnaire (fr) · Regards sur l’actualité de la Wikimedia (fr) · Wikimag (fr) · Education · GLAM · Milestones · Wikidata · Central and Eastern Europe · other newsletters

Subscribe or unsubscribe

email encryption

Email, one of the most pervasive and enduring technologies of computer networking, was invented in about a dozen places by dozens of people in the 1960s. It's hard to lay out a clear history of the technology because it's just so obvious—pretty much as soon as more than one person could use a computer, there was some kind of mail facility. These ranged from mainframe-centric systems where all of the users of a single computer could write messages to each other, to PC-centric systems where workstations would mount a network share to store and retrieve messages. Pretty much any scheme you can come up with for moving messages was probably in use somewhere from roughly the 1960s to the 1990s, by which time the ARPANET-derived family of email implementations had taken hold.

This form of email has a clearer heritage, to Ray Tomlinson, who came up with the core idea that addresses could identify both a user and a host, and that some kind of open protocol could be used to send messages to another host when necessary. Over time, and with many revisions, Tomlinson's design became SMTP and was joined by protocols like IMAP that built out the form of email we use today. This is a form of email that is in some ways decentralized (any user is free to choose a host) and in other ways centralized (each host assumed to be continuously online to store-and-forward messages for its users). Tomlinson's design was flexible enough that we have not had to totally get rid of it, but enough has changed about the modern Internet that we have had to take a new approach.

Email has many vexing limitations, artifacts of its age. For example, email handling should not be assumed to be "8-bit clean"—email protocols were originally defined over 7-bit ASCII and ran on many machines that used the eighth bit as a checksum. These machines were prone to changing the last bit of each byte, or otherwise mishandling email with 8-bit content. That wasn't a problem when text was completely limited to that 7-bit plane, but both Unicode and the desire to send binary files made 7-bit email unworkable. MIME was developed as a workaround, an encoding technique that solves a few problems in one go by encoding all non-ASCII content of email in the form of ASCII characters.

This is rather inelegant. MIME's ASCII encoding, whether quoted-printable style or base64 style, is annoying for humans to read and inefficient with regard to message size. As with all base64-like encodings, it counts as a suspicious smell, an indication that we are papering over past mistakes. Email is full of them. Let's consider another: security.

"Internet email," a term we can use to try to explicitly specify the SMTP-type email system inherited from ARPANET, uses a simple design that dates to an era when the operators of networked computers all knew each other. While there were other possibilities, most email was delivered directly to the computer of the recipient, and in a lot of cases the host-level security of the era wasn't very good either. The possibility that someone else could read your private email was known and there wasn't a whole lot you could do about it. The internet existed within a context of trust.

We don't extend that same trust today. Email is usually handled by third-party mail services on either end, which are not necessarily trustworthy, and to get between the two third-party mail services it has to pass through a series of internet links that are also of questionable security. Email became one of the first, probably the first, compelling application of large-scale computer networking. It may have also been the first of the field's enduring security failures. The system was designed with no concern for message confidentiality or integrity, and much of it still operates that way today.

Still, there's a long history of efforts to improve email security, and we should probably start the story with the greatest of all: not just an email protocol, but an entire stack of protocols that promised a better internet in general: the OSI suite.

X.400

The modern internet was born of an era often called the "protocol wars." When we tell the story of the internet today, we tend to simplify it to a nice linear tale in which ARPANET was created, everyone thought it was a good idea, and it proliferated. That's not totally wrong, but it leaves out a lot of the other actors. The telecom industry famously had its objections to the packet-switched model, which had originated from the computer industry rather than the communications industry and showed it. During the 1970s and 1980s, the telecom industry defined its own set of network protocols, and fundamental approach to computer networking, which came to be known as the OSI protocol suite.

Today, the relevance of the OSI protocols is limited to the frustrating insistence of modern networking textbooks on trying to explain the internet in terms of the "OSI model," a now more speculative than actual layer model that describes a different system incompatible with the internet. The reason that academia has this weird fixation on using the design documentation of a failed internet to explain the fundamentally different architecture of the other internet that succeeded... well, that itself is fallout of the protocol wars. History may be written by the victors, but the losers still get to dictate theory.

Despite my ongoing project to discourage the use of the OSI model for learning, we can indeed learn from the OSI protocols. The reasons that the TCP/IP stack succeeded over the OSI stack are many, one of them being that OSI's "design by committee" approach led to a set of protocols that are very thoroughly thought out but quite complex. This is a big difference from the internet protocols, whose design philosophy was closer to "good enough for now."

The difference in philosophy is especially stark when we look at the application-layer protocols. The "internet protocol suite" as it is conventionally defined does not include application protocols at all, itself a telling fact. The OSI suite does. Internet email started out as a special-case use of FTP, before moving over to a loose set of commands over Telnet that eventually formalized (or ossified, depending on who you ask) into SMTP.

By the time the OSI standards were published in 1984, it was known that email was one of the killer applications—the first revision of SMTP was published in 1982, a post-facto response to the popularity of email on the internet, and besides, email had been a key feature of most proprietary networking systems since the 1970s. So, OSI had a solution: X.400.

X.400 is the "Message Handling System" part of the OSI suite, and provides features which are recognizable as email but also a bit more generalized. X.400 was, in most regards, a failure. It was unsuccessful in displacing internet email in almost all environments. Still, X.400 was decidedly influential. There was a period, albeit brief, in which it was thought that major governments would require X.400 compatibility in future purchasing. This was similar, and had a similar effect, to the period where POSIX-compatibility was a government requirement: a lot of vendors designed their products to meet the requirement, the government lost interest, and it became a curious legacy. In email, this is most obvious in the form of Microsoft Exchange: Exchange started out as an X.400 implementation, mostly for government reasons, and still provides a lot of X.400 features.

X.400 is vastly more complex than internet email. For example, the encoding of X.400 messages is ASN.1, a binary serialization format that is more efficient and capable than MIME but also a lot more complicated. That X.400 has complexity stacked on top of complexity is useful in understanding why one of X.400's most compelling features didn't go very far: encryption.

If you are sending an email to someone, and you don't want anyone else to read it, the most obvious approach is end-to-end encryption. For many practical reasons, you'll want to use asymmetric encryption. You find a public key for the recipient, encrypt the message to that key, and then send it over the wire as usual. The recipient then decrypts it with their private key. Simple enough, except that every part of this process is laden with complexity.

To underscore that point, let's take a look at X.400 itself—the 1988 "Blue Book" version, which seems most appropriate to when email was ossifying.

Aspects of an asymmetric key management scheme to support the above features are provided by the directory system authentication framework, described in Recommendation X.509. The directory stores certified copies of public keys for MHS users which can be used to provide authentication and to facilitate key exchange for use in data confidentiality and data integrity mechanisms. The certificates can be read from the directory using the directory access protocol described in Recommendation X.519.

Recommendations for other types of key management schemes, including symmetric encryption, to support the security features are for further study.

End-to-end encrypted email requires two basic parts:

First, there must be a technical format for moving encrypted messages through the message transfer system. X.400 handily resolves this by the use of ASN.1 definitions that allow for encrypted messages bodies, thus making most of that someone else's problem. X.400 actually contemplated a lot more when it comes to security features, things like non-repudiation of delivery, that required more complexity in message transfer but didn't make it to the modern day. We'll just ignore those for now.

Second, there has to be a key infrastructure. You need some way of obtaining a public key for the person you want to send a message, and of being sure that it really belongs to them. This is important both for encrypting outbound email and for authenticating inbound email, since end-to-end encryption is usually deployed alongside end-to-end authentication using the same cryptographic infrastructure. In the case of X.400, that whole problem is deferred to the X.500 family of protocols.

What is X.500? It's our old friend, directory access protocol. The X.500 directory standards are the comparison by which LDAP is "lightweight," and there is a parallel story: LDAP is the scrappy internet protocol that won against the X.500 Goliath.

So the gist of X.400 email security is this: use the directory to look up the recipient's public key, and then encrypt the body of the ASN.1 message object to that key. In actuality, the details are different and more complicated, but that goes for everything OSI and the simplified sketch here is enough to get the gist. It's straightforward.

There are just two problems: ASN.1 as a format for messages did not succeed, and the concept of a global directory never materialized.

To the extent that X.400 succeeded, its security features were actually a big motivation. The ISODE Consortium, which had created something like a reference implementation of much of the OSI application suite, is still around today as Isode Ltd and still sells a message handling system. There are a few other maintained X.400 descendants as well. One of the largest customer bases is military: while less common in the US (keep in mind that internet email was originally developed under the aegis of the US military), a number of European countries adopted X.400 as their secure messaging solution and still have extensive reliance on it in military and intelligence organizations. Other enduring strongholds of X.400 include global aviation (ICAO's messaging system is X.400-based) and electronic data interchange, or EDI, a standardized system of financial and supply chain messaging between ERPs.

It is no coincidence that these are the same kinds of applications where protocols like X.25 survived unusually late (indicating a high affinity to long-running, legacy systems) and that they are contexts where the participants and messaging are governed by a central entity. That means that a directory exists, which means that the key distribution problem is vastly simplified. It also means that the software in use, both transfer agents (e.g. servers) and clients, are standardized and usually purchased from vendors that have specifically designed them to work well in an X.400 environment.

In other words, they are contexts that are very different from email, in which people exchange messages between a huge number of organizations and many users are simply on "mail providers" that are not an organization in a traditional sense. This means that there is no meaningful directory capability in the email system. Email is also used with a hugely diverse landscape of clients, imposing a steep penalty on the "complicated message encoding that requires interaction with a complicated directory system" part of the design.

Since X.400 proved a nonstarter in most of the email world, what about encryption for the rest of us? Well, this situation gets complicated and sad.

MIME

First, it's helpful to expand a bit on the involvement of MIME, the Multipurpose Internet Mail Extensions. MIME originated as a solution to the 8-bit problem with internet email, as a way to encode non-ASCII text and binary files in a consistent, reliable way. Along the way to achieving this aim, it introduces the concept of a "multipart" message which has multiple distinct objects as a body instead of just one string of text. We use this multipart feature heavily today, both for obvious applications like file attachments and for more subtle ones like HTML emails that provide a plaintext variant. I don't think it's totally wrong to summarize MIME as "email 2.0," because while it is limited in scope to the message format (not changing the transport protocols), MIME adds a lot of new functionality to email that we now consider basic requirements for a messaging system.

The history of MIME has a lot more twists and turns than you might expect, which reflects the difficulty of making major changes to a system as distributed and heterogeneous as internet email. Still, the MIME authors seem to have had a similar view that MIME was a significant evolution of email by standardizing message bodies and paving the way for future extensions to that standard, adding new functionality. This quote from RFC 1521 (1993), the first MIME standard, is informative:

STD 11, RFC 822 defines a message representation protocol which specifies considerable detail about message headers, but which leaves the message content, or message body, as flat ASCII text. This document redefines the format of message bodies to allow multi-part textual and non-textual message bodies to be represented and exchanged without loss of information. This is based on earlier work documented in RFC 934 and STD 11, RFC 1049, but extends and revises that work. Because RFC 822 said so little about message bodies, this document is largely orthogonal to (rather than a revision of) RFC 822.

In other words, existing email standards addressed how messages are conveyed between hosts, but did little to address what would actually go into those messages. MIME addresses that problem by creating a proper standard for message bodies, one that adds features while contending with the unfortunate limitations of the many pre-MIME email implementations. The many RFCs around the general topic of MIME (and other proposals for email body formats) spend a lot of time contemplating how these new message types would be handled by email clients that did not understand them. Email technology is generally easy, email standards are hard.

PGP

As these standards efforts marched on, the fast-growing field of computer cryptography had its own ideas. The first major event was cryptographer Phil Zimmerman's unceremonious 1991 release of a tool called PGP, or Pretty Good Privacy. The story of PGP is a long one, jumping from a casual release to an activist Usenet group to criminal charges to the foundation of PGP Inc to commercialize the system.

The funny thing is that PGP itself is not that important to the story of PGP. Despite the corporate ambitions of the company bearing that name, PGP started as a loosely academic or activist project, and it ended up that way in the long term as the open-source clone GnuPG (GPG) displaced PGP in the marketplace. PGP Inc became part of Symantec and faded into the obscurity that most Symantec products do. Along the way, PGP and GPG implementations rode the roller coaster that was 1990s cryptography, swerving from algorithm to algorithm and growing a complex cipher selection system along the way.

PGP was a general encryption tool and could be used for files or messages of any kind, but email was a key application. So how is a PGP-encrypted email actually conveyed? Well, it's not exactly pretty, but it aligns with the other messes we've already seen. PGP typically outputs ciphertext as either binary data or in a format called "ASCII-armored," which is base64 except with an optional CRC checksum trailer (which was deprecated decades ago). This is very similar to the PEM format widely used with modern cryptographic applications, and not by coincidence: PEM stands for Privacy-Enhanced Mail and is itself an artifact of a failed email encryption standard that produced quite a few RFCs. PGP's email application and PEM were under development just about simultaneously, and there was definitely cross-pollination of ideas between the two. PGP won, but not without taking parts of PEM along for the ride.

In practice, ASCII-armored PGP payloads are sent one of two ways: first is "inline" format in which each part of a MIME-encoded email is separately encrypted and have otherwise usual MIME types (e.g. application/octet-stream), so that non-PGP-capable email clients will treat them as opaque files that the user can save off and process with a separate tool. The second approach is the "MIME" format in which the entire MIME-encoded email is encrypted whole and dropped into a new MIME document with the type "application/pgp-encrypted." The latter has better security properties (less metadata leakage) and tends to be more user-friendly and elegant in clients that are designed with PGP in mind, but fails very badly in non-PGP-capable clients which are not even able to separate out attachments. You can still process the email with an external PGP tool by saving off the whole body, MIME-decoding, decrypting, and MIME-decoding again, but you get a little frustrated just reading that, don't you?

Today, this mess has become one of the two main criticisms of PGP. PGP is not the result of a huge, multi-party standards effort like X.400, but rather the product of a small group of people taking a pragmatic approach. Nonetheless, it grew up in the same climate and so survived with most of the same scars. From bottom to top, PGP is very complicated. Almost every part of PGP-encrypted email has multiple options, some of which are known to be insecure. Modern tools mostly follow sane defaults, but plenty of older tools remain in use. Key tools like gpg itself show their age with some of the least user-friendly interfaces ever designed. The immense difficulty faced by PGP users was the subject of seminal 1999 "usable security" paper "Why Johnny Can't Encrypt" (the paper that more or less invented the field of "usable security"), and most of the criticisms raised in that paper are still true of PGP implementations 27 years later.

And that brings us to the second aspect of PGP, key distribution. We've seen, at least in sketch form, how PGP email is conveyed, the first challenge of encrypted email. But what of the second challenge—is PGP for use with a directory system? Well, the earliest examples of PGP de facto were, because they did not address key distribution at all. Very early on, though, PGP gained a novel concept called a "web of trust." The web of trust, or WOT, is simultaneously one of computer security's most intriguing and appealing concepts and one of its most inane and stupid. I'll try to explain it succinctly.

You are Alice, and you want to send an email to Bob. The "why" is not important, these are timeless characters of cryptographic explanations and they have a timeless need to securely exchange messages. If you know Bob in person, say because you are colleagues at an institution, the key distribution problem is simple. You can perform calculations on a key, like a checksum, to produce a "fingerprint" which is unique to the key and short enough to not be completely impractical to read by eye. You, Alice, walk over to Bob's office and ask to see the fingerprint of his key. You write that down on a sticky note, and then in the future, you make sure to encrypt your emails to Bob using a public key that has that correct fingerprint. Easy as pie so far, but that doesn't scale well to situations where Bob is not within walking distance.

And so the "web" enters: Bob may not be conveniently accessible to Alice because, say, he works somewhere else, but Alice's coworker Charlie might have been at a conference with Bob a few months ago. If Charlie took down Bob's key fingerprint, he can give Alice a copy when he gets back. Since Alice knows Charlie and believes him to be trustworthy, this is a pretty good substitute for getting the fingerprint straight from the horse's mouth.

PGP's WOT model encodes this sneakernet approach to key distribution into a software implementation. A PGP user can "sign" another user's key, a way of cryptographically attesting that they believe that to be the correct key. These signatures are uploaded to public "keyservers" that allow anyone to query for a key and for the set of signatures it has received. If there are enough keys with enough signatures, you can treat the whole thing as a graph and try to find "trust paths" between Alice and Bob, even if doing so requires six degrees of separation1. The tangled interconnection of keys and signatures verifying them forms a web: the web of trust.

In practice, the tide seems to have turned against the WOT model by the 2010s, and today it has been completely abandoned. It didn't scale as well as its proponents had hoped, usability problems with the tooling discouraged participation, and both the implementation and concept were complex enough that poor user understanding undermined the benefits. Today the WOT has fundamental problems (mostly related to the open keyservers) that have made it practically unusable, and I think it is fair to call it a failed experiment. Further extending that fairness, we should note that the WOT was surely an inspiration for the vastly simplified key verification schemes that appear in most modern end-to-end encrypted messengers.

From its messy academic origins to its transformation into an open project, PGP enjoyed a lot of prominence in academic and "cypherpunk" communities during the 1990s and 2000s, but was less common in corporate environments. The WOT model only made sense in the absence of a centrally maintained directory, and while PGP could be used with a directory it was not the path espoused by its advocates. PGP was mostly used among researchers and hobbyists, the cryptographic establishment. In the Enterprise, there was another player: S/MIME.

S/MIME

S/MIME originated with RSA Security, the company founded to commercialize the RSA encryption algorithm. RSA Security survives today as a subsidiary of teacher's pension funds, although its prominence in the industry has generally declined. Back in the 1990s, RSA was on the cutting edge, so 1996's introduction of an RSA solution for email encryption got quick attention from industry.

Over the following two years, RSA's product evolved into an IETF standard called S/MIME, for Secure MIME. S/MIME saw a series of new versions up to the present day, and is exceptionally rare in the academic and hobbyist world but fairly common in corporate and government environments, especially with Microsoft Exchange (having adopted S/MIME as its secure email option of choice after the failure of X.400 to take over inter-organization mail). Let's see how S/MIME handles the first challenge of encrypted email, encoding the message for transport.

Well, as the name suggests, MIME is involved. The simple sketch is like this: an email is prepared as a MIME document, which is then encrypted into the form of a PKCS#7 document. PKCS, the Public Key Cryptography Standards, refer to a series of open standards memos published by RSA Security. PKCS#7 describes "Cryptographic Message Syntax" or CMS, a standard container for an encrypted payload along with various metadata. In the terminology of the time, S/MIME is said to "enhance" messages with the "security services" of PKCS#7. We still use the term PKCS#7 today, mostly to refer to cryptographic certificates serialized in that format. For email, on the other hand, PKCS#7 has been superseded by an IETF standard called simply "CMS."

The CMS version of the email, which remember is an encrypted MIME document plus cryptosystem-related metadata, is then encoded as PEM (functionally Base64) and made the body of a MIME email with a type of "application/pkcs7-mime" (yes, I just said that PKCS#7 is gone from this application, but naturally the name has stayed). This is pretty much equivalent to the MIME version of PGP emails, and has basically the same upsides and downsides: it's relatively elegant and avoids a lot of weird edge cases in clients that support it, but it's quite hostile to clients without S/MIME support, which will be unable to make anything of the message. It also keeps many of the general downsides of MIME, and for example the RFC has to spend some paragraphs justifying why binary data is base64d and then converted to binary and then base64d again. Not having the extra layers of base64 involved would break interoperability scenarios with non-S/MIME tools!

Overall, S/MIME isn't that dissimilar from PGP in this regard, although I would say that it is overall simpler and better standardized. That is not to say that it is "simple" or "well-standardized," these are only relative observations. S/MIME implementations are still quite complex, and interoperability problems can appear even between organizations running the same software, due to the many configuration knobs involved. RFC 2311 includes this gem, to give some of the flavor of it:

S/MIME provides one format for enveloped-only data, several formats for signed-only data, and several formats for signed and enveloped data. Several formats are required to accommodate several environments, in particular for signed messages. The criteria for choosing among these formats are also described.

"Several formats are required to accommodate several environments" is the story of email in a nutshell... if not all of computing.

Let's consider the second aspect of the encrypted email problem, then: key distribution. Here, S/MIME takes a very X.400 style dodge. The original S/MIME standards documents, as far as I can tell, do not address key distribution in any way except for oblique and incidental references like "validating the certificate to a trust anchor." This seems odd except in the context of S/MIME's development by RSA Security and, as it became an IETF standard, Netscape—the birthplace of SSL. S/MIME was intended for use with the PKI, the Public Key Infrastructure, which has a nuanced and sometimes vague history but fell out of a combination of X.500 concepts and implementation work by companies including RSA Security and Netscape. In other words, S/MIME grew up alongside SSL/TLS, and uses the same key distribution architecture.

S/MIME emails are expected to provide the certificates that go along with the keys used (one of the reasons for the CMS container format), and those certificates are supposed to be signed by a Certificate Authority.


I put a lot of time into writing this, and I hope that you enjoy reading it. If you can spare a few dollars, consider supporting me on ko-fi. You'll receive an occasional extra, subscribers-only post, and defray the costs of providing artisanal, hand-built world wide web directly from Albuquerque, New Mexico.


This is not exactly the same as saying that keys will be retrieved from a directory, but if you remember that the X.500 that gave us these concepts was itself supposed to be the directory, you can see that they are closely related ideas. In practice, they are often pretty much the same: a classic real-world application of S/MIME is in a Microsoft Active Directory environment, where the Domain Controller acts both as a directory (LDAP) and as a Certificate Authority, signing certificates for the users that populate that directory. This means that key retrieval from LDAP vs. certificate verification against the (AD) CA are pretty much the same thing, and you have the same assumptions around there being some sort of central authority to maintain the directory.

This means that S/MIME thrives in the same places that X.400 was more successful. The US military, birthplace of internet email, opts for S/MIME for encrypting it. They have a truly enormous Exchange and Active Directory environment to this end. Once those emails leave the military, the whole thing becomes a huge hassle, because properly using S/MIME requires establishing some form of federation with the military's Active Directory domain... or at least importing their CA as trusted, which is not all that easy with most mail clients.

Today

So here we have the classical email encryption landscape: there are two major options, PGP and S/MIME, each of which is unsuitable for most users for its own special reasons. Out of this frustrating stalemate, several other phenomenon have emerged.

First, we should briefly touch on "modern" email encryption, where "modern" can become a bit of a euphemism for "proprietary." There is a new set of email providers that offer some kind of built-in, managed E2E. Protonmail is the most prominent. Protonmail's dominant E2E encryption implementation is a proprietary, simplified system that relies on Protonmail as key directory and cryptographic implementation. This is not interoperable with other email providers, so I do not consider it a true form of "email encryption," but rather a proprietary encrypted messaging product offered alongside email. Protonmail does have two features for sending secure email to non-Protonmail recipients: the first is essentially a compatibility layer to PGP that allows Protonmail's encryption to get along with PGP tools. The second is a traditional compliance-driven "secure email" product where the recipient is actually sent a link to a website where they can enter a password to retrieve the message—once again, not really email but a separate product that loosely works with email.

Second, let's talk a bit about how email is really encrypted, in real reality, because there is this huge, grand-canyon-esque disconnect between the way that people discuss email encryption and what actually happens in the real world. Let's start taking that on by differentiating some terms: so far I have discussed end-to-end or E2E encryption, which is not exactly a well-defined technical term but generally means that a single encrypted payload travels unmodified and undecrypted from a human user to a human user. This can be distinguished from non-E2E encryption, where a message might be encrypted and decrypted several times, and using keys controlled by people other than the sender and recipient. We should also note the distinction between "encryption at rest," whether or not data is encrypted when stored (e.g. on disk), and "encryption in transit," or whether or not data is encrypted when sent over the network.

Email is exchanged between email servers using SMTP, and is exchanged between email servers and clients using IMAP 2. Both of these protocols are now widely used over encryption, either using the STARTTLS method or "real" TLS, a distinction that I will not explain right now. While I lack solid data, it seems like a safe bet that the vast majority of real-world email is encrypted in transit, because at least the mail server logs that I have access to show the vast majority of mail servers using STARTTLS to secure their transfer connections. Client configuration is a more confusing topic but everything other than IMAP over TLS is now generally regarded as a big mistake, and modern clients should aggressively steer the user towards that route.

This is a confusing contradiction: we have had this big, long, boring discussion of how email encryption has failed to go anywhere. Then I tell you that almost all email is encrypted. What gives?

To start, remember our "in transit vs. at rest, and E2E vs. not-E2E, distinctions. The fact that the network connections between email participants are generally encrypted gives us a lot of security advantages, but it does not address the full scope of requirements that PGP and S/MIME do. Your email is not secure from your email provider, or the recipient's email provider for that matter, since email is usually not encrypted at rest and the provider has all the keys anyway. It also does not provide integrity protection, or non-repudiation, or any of a number of other useful security properties.

What's even worse is the technical reality. I pointed out that one of the big criticisms of PGP is that it is simply too complicated and provides too many options. This resulted in a litany of defects in PGP implementations over the years, and no small amount of insecure user behavior due to lack of usability. In-transit email encryption has very similar problems. STARTTLS, for example, is a fundamentally opportunistic mechanism. In its original form, an on-path attacker attempting to intercept an SMTP session can simply hijack it and suppress the STARTTLS offers from the two sides, preventing them ever negotiating a switch to encryption. This simple "downgrade attack" is a pervasive problem, and it's the main reason that I say that anything other than true TLS is a mistake. Unfortunately, STARTTLS is the most widespread option available for inter-provider email transfer.

That said, there are ways to address this problem. For example, DANE, the DNS-based scheme for TLS certificate distribution, has failed in the HTTP arena but has decent adoption (although very far from universal) in email transfer. It does a good job of providing mail servers with a way to discover that their peers are capable of encryption, without doing so over the SMTP channel itself. Even without DNSSEC to fully protect the DANE channel, this is an improvement that prevents many common attack scenarios. There is an even more popular approach that, while clumsy, avoids the dependence on DNSSEC to avoid a second on-path attack on the key channel: MTA-STS. MTA-STS is similar to DANE in that it provides an independent channel for a mail server to find if a peer supports TLS, but it uses HTTPS for that second channel, thus inheriting the protections of CA-rooted TLS.

Both DANE and MTA-STS are ways of "tipping" a mail server that a peer has STARTTLS support, at which point the mail server knows that it should only connect via STARTTLS, even if the mail server on the actual SMTP channel claims not to support it. While the extra moving parts are unfortunate, this does a lot to close the downgrade attack window, and email transferred between mail providers with DANE or MTA-STS support is really pretty well secured, at least on par with most instant messaging solutions.

This also implies that you can improve your email security a lot by just manually pinning major peers—e.g. configuring your mail server that, regardless of everything else, connections to major providers like gmail.com must be encrypted. This is actually a pretty common practice in some industries, including internal processes to maintain the list of "secure email destinations" so that it includes business partners and other frequent recipients.

In some industries, this is taken to an extreme: within the healthcare sector, for example, there are healthcare information exchanges that operate secure messaging platforms that are mostly just conventional email servers configured to refuse any non-TLS SMTP exchanges. This is sufficient to meet NIST 800-43 recommendations which is sufficient to meet the NIST security architecture for healthcare information. That might be surprising to you but it illustrates how this topic gets confusing: we all know that "email is not secure," but that's really a configuration problem with the "public" email system rather than a technical limitation. Few compliance standards actually require E2E, and securing email in transit and at rest, once you drop the E2E requirement, is so easy that it sometimes happens just by default when SMTP servers have STARTTLS support and disk encryption.

If you look at healthcare marketing materials, you'll sometimes see STARTTLS and S/MIME described as alternatives to each other. This doesn't make much technical sense, but you have to think about the context. If you have to communicate over mail servers whose configuration you do not control (or at least have insight into), you don't have a way of ensuring that your messages will be encrypted in transit unless you use E2E encryption such as S/MIME. If you are sending mail between servers that you do control, though, you can ensure encryption in transit by configuring the mail servers that way, and you do not need E2E encryption to meet that specific aim. Taken in this view, E2E encryption is almost a workaround for a defect of the design of email, and if you have taken anything from this longer than expected journey I hope it will be that: the history of encrypted email is so odd, and full of so many false starts, because it is basically the history of E2E encryption as an overlay "fix" for email's lack of the level of security provided by newer messaging solutions out of the box.

This also helps to explain the lack of adoption: it's not really that we can't figure out how to encrypt email, we do encrypt email. But as with most communications technology, the "security layer" has been pushed onto the service providers rather than the users. This has significant downsides, since it requires trust in those service providers, but one huge, conspicuous upside: if encryption is transparent to the user, people will actually use it.

  1. The idea that all humans are separated by six degrees at most was popularized in 1990 and undoubtedly had influence on the PGP WOT design, as well as a number of other similar connection-following social technology ideas of the same period. As it turns out, the whole idea came from fiction and has little grounding in actual statistical research on human social networks, which is at least suggestive of a reason why none of these web models have been successful.

  2. I am intentionally simplifying this situation here, both by addressing only the case of IMAP and using the terms "server" and "client" for email, which only loosely map to the more proper roles of "transfer agent" and "user agent."

megawatts by microwave

In 1914 the Department of the Interior, through the Bureau of Reclamation, investigated the possibilities of developing the Columbia River. Thousands of arid but potentially fertile acres needed only water to become the Imperial Valley of the Northwest. Locked in the mountain ranges were valuable ores awaiting electricity to turn them into needed metals.

Two years later the State engineer of Oregon urged the development of the Bonneville site as a national-defense measure: he saw in the proposed power project a source of fertilizer in time of peace and nitrates in time of war. The dam also would completely drown out the Cascade Rapids and extend slack-water navigation some 40 miles eastward to The Dalles.

The Rivers and Harbors Act of 1925 directed the Secretary of War, through the Corps of Engineers, United States Army, to prepare and submit to the Congress an estimate of the cost of surveys, examinations, and investigations of all navigable streams and their tributaries where power development appeared feasible. (Q1)

It is difficult to succinctly explain why, exactly, the United States Army has spent much of its history involved in the construction of dams. It is partly an accident of history, partly the result of interagency federal politics, and entirely a product of American culture. In his book "Cadillac Desert," Marc Reisner examines the history of the American West's water control projects as a religious project, one animated less by practical needs than by a sense that domination of the West's rivers was destiny.

The Bureau of Reclamation, part of the Department of the Interior, was formed for that purpose. At the time, though, the Army had already been used to survey and improve rivers for nearly 100 years. They were not content to give it up. The result was a rivalry, one with several feints and blows before the two settled into their modern areas of control. For the Bureau of Reclamation, the Hoover Dam was their signature project. For the Corps of Engineers, the battle that would go down in history was the Columbia River Project.

The motivations for damming the Columbia were various. The Columbia was prone to flooding, which had caused damage and limited use of land along it. There was a great deal of land surrounding the Columbia that could be farmed, if the Columbia could be tapped for irrigation. Electricity, too, was a reason, although initially a somewhat secondary one. Perhaps the greatest reason, though, was simply economic: by the time that the major parts of the Columbia River Project were truly underway, the nation was in the throes of the Great Depression.

President Franklin D. Roosevelt was already a fan of hydroelectricity. As governor of New York, he was exposed to the pioneering Niagara Falls power plant and pushed for other similar projects in that state. As President, his "New Deal" naturally incorporated hydropower as well. By 1934, he had formed a Regional Planning Commission that sketched out a series of dams along the Columbia, two of which would become the Grand Coulee and the Bonneville. These dams would produce a tremendous amount of electricity, and unlike in other similar Corps of Engineers projects to date, that power would not all be consumed by irrigation pumping. There was power to spare. To distribute that power, the Regional Planning Commission suggested an independent government agency on the model of the Panama Canal or the recently chartered Tennessee Valley Authority.

As an interim measure, the loosely defined Bonneville Project coordinated the civilian side of the Corps of Engineers project until 1938, when the Bonneville Dam was complete and the Grand Coulee was much of the way there. The Bonneville Dam captures little water in its reservoir, so while it does have flood control value, electrical production is its primary purpose. The dam's two powerhouses produce up to 1.2 GW, an impressive number for the 1930s but one that pales in comparison to the Grand Coulee's eventual (1970s) full capacity of nearly 7 GW. The Columbia River dams increased the electrical capacity of the Pacific Northwest by orders of magnitude; the numbers were significant even at a nationwide scale.

The bumper crop of electricity triggered a predictable controversy: what to do with government power? One camp favored public control of the resource, with the government marketing the power on some sort of equitable basis. The other favored private control, arguing that the output of the dams should be contracted entirely to private utilities like Portland General Electric (itself the scion of an important early hydroelectric project at Willamette Falls). In the New Deal political climate, the first camp won: the Columbia did not quite get a TVA, but Congress did charter the Bonneville Power Administration (BPA), the first of what would come to be known as Power Marketing Agencies. Over the following decades, the BPA became part of the Department of Energy (DoE)—uncharacteristically, for the DoE, a part of it that actually generated and sold electricity. Well, technically, the Corps of Engineers generates it, and the BPA markets and distributes it.

In any case, starting in the late 1930s, the BPA was tasked with the construction of a network that could distribute power from Columbia River dams throughout the region—on an equitable, equal-rate basis often called the "postage stamp rate" that allowed rural coops to buy government-generated power at the same rate as the big city private utilities. The sudden bevy of power along the Columbia and the fair rates at which it could be obtained in great quantity led to an industrial revolution for the region, one that saw it as the seat of the American aluminum industry (with the Columbia Gorge producing something like 1/3rd of the nation's aluminum through to the 1970s) and that boosted the fate of hundreds of related industries (aerospace and, specifically, Boeing not least among them). BPA power has enduring influence today, with many towns on the Gorge (The Dalles, Boardman, Umatilla) disproportionately prominent on a map of the nation's data centers. AWS's us-west-2, for example, is a beneficiary of Columbia River dams and located near many of them—not just Bonneville, but the Dalles (1.8 GW), John Day (2.2 GW), McNary (1.1 GW), and more.

In marketing this power, the BPA faced a challenge: the dams are spread across a large area, as are the customers. Industrial customers, such as the Alcoa (Aluminum Company of America) smelter that opened in 1940 at Vancouver, Washington 1 were opening in rural areas where land was readily available, and an explicit goal of the Columbia River Project had been the extension of electricity to farmers and other rural industries. The concept of long-distance power transmission had been pioneered by an 1889 transmission line, the nation's first, between Willamette Falls at Oregon City and downtown Portland. Beginning in 1938, the BPA was tasked with expanding that concept across a region that would eventually span eight states.

The Master Grid

BPA's first administrator, J. D. Ross, presented a plan he called the BPA Master Grid. This ring-shaped network, made up of 230 kV long-distance transmission lines, would connect the dams not only to Portland and Seattle but to Pasco, Yakima, Spokane, Ellensburg, the Willamette Valley through to California, and the Oregon Coast. By 1945, the Master Grid covered three thousand "Circuit Miles" of transmission lines. It was the first integrated regional power grid in the United States, and would come to pioneer the market-based electricity pricing and distribution, independent system operators (ISOs), and pooling and wheeling agreements that form the modern US electrical infrastructure. The entire Western Interconnection, the unified power system that serves the US and Canada from the Rocky Mountains west, can be said to have crystallized outwards from the seed of the Bonneville Dam's switch yard.

Getting there required that the BPA solve formidable technical problems and develop many new technologies in power distribution. BPA transmission lines operated at higher voltages than any before them and, in the 1960s, introduced high voltage DC transmission to the Americas, connecting the Columbia system to the major demand centers of Southern California at 800 kV DC. BPA was only slightly behind the TVA on the installation of a remarkable analog computer called a Network Analyzer, in 1939, which simulated the behavior of the transmission network like a scale model. The rural nature of the BPA network put substations in remote areas, where they were minimally staffed, and the long stretches of high-voltage transmission line meant there was ample potential for damage by wind, weather, and trees, phenomena that the BPA came to better understand through research laboratories and experimental field sites.

This is not an article about the history of electrical distribution, or at least it wasn't supposed to be, so here we must exercise some discipline and narrow in on a topic. Telecommunications ought to do.

By 1940, as the Master Grid entered operation, its numerous substations already caused administrators a headache. Each had a small staff of technicians, but communicating with them was difficult. Coordinating changes across large areas, or quickly responding to faults, involved a flurry of telephone and radio calls. When Portland General Electric built the transmission line from Willamette Falls to Portland, they encountered the same problem, and by the 1910s had implemented a very early form of its solution: telemetry and teleoperation. Through a set of control wires strung along the transmission line, operators in Portland could see certain measurements from the Oregon City powerhouse and remotely throw switches to bring turbines on and offline in response to load. As the BPA built the Master Grid, they invested in the same technology.

Around 1939, the BPA commissioned a study of communications technology that could be used along the Master Grid. There were three main contenders: commercial telephone networks (which BPA called "land telephone" to differentiate it from the other two), "carrier current telephone" technology that superimposed telephone signals onto the electrical conductors of the transmission lines themselves, and radiotelephone equipment. A working agreement was reached with Pacific Telephone & Telegraph, the Bell System company that would later become US West, to share network information and analyze the cost tradeoffs between purchasing carrier current and radio equipment and leasing telephone lines. Ultimately, the diversity of the BPA network required some of all three.

Each of the BPA's substations had a building, called the control house, that contained control and monitoring equipment along with office facilities for the substation's operators. A room of each control house was dedicated to carrier equipment, devices that modulated multiple telephone circuits using frequency division multiplexing, and to a set of carrier frequencies that could be coupled onto the transmission lines to be received at the next substation. This equipment is similar to carrier equipment used in the telephone network, although specialized to power distribution applications by the choices of carrier frequency. I cannot say for certain, but it is very likely that BPA purchased their system from Lenkurt, a San Francisco-based communications equipment manufacturer that specialized in carrier current systems at the time 2.

The BPA's carrier system incorporated selective calling, meaning that users interacted with telephones that looked and felt much like conventional telephones, including a dial. An operator at one substation could dial the number for another and that phone would ring. The main difference from the telephones we use today is that these carrier current systems were interphones, more similar to intercoms or party lines than single-user telephone service. If you picked up a phone on a circuit, you would hear any conversations already underway. Of course, in industrial control applications, this built-in conferencing capability was generally considered a feature, and telephone circuits were assigned to shared use by departments or operating regions.

Radio was installed as well, primarily so that construction and maintenance crews in the field could get messages back to the administrative offices. Some substations, in strategic locations for coverage, had a radio site about a half mile from the substation for isolation from the powerful electromagnetic interference created by the high-voltage transformers. These radio sites were wired to remote heads located in the substation control house office, where substation operators relayed messages between mobile radios in the field and the carrier current telephone system. Bear in mind that these were still early days for mobile radios, and the HF units used by the BPA were proudly described as using only 1.5 cubic feet of space in the trunk of the vehicle, plus the microphone, speaker, and control head in the cab. Finally, while not the main purpose, it was already noted in the 1940 Annual Report that the substations could use the radio stations to substitute for the carrier current system in an emergency.

Complementing all of the above, the BPA leased telephone lines between major substations, agency headquarters in Portland, and the switch yard in Vancouver that was becoming the closest thing to a "main substation" in the network's distributed, ring-shaped design.

Immediately after the BPA's first Annual Report discusses the selection of communications equipment, it moves on to Protection. Here I must introduce a topic in electrical engineering that I have only a loose understanding of, even having spent the last few weeks in part on YouTube engineering tutorials. It's important that we get comfortable with the field of "protective relaying" because, as we will see, it became the most widespread application of private telecommunications networks after the railroads.

Protective Relaying

In your home, you are protected against certain dangerous scenarios by the over-current protection device that we Americans call a circuit breaker. The electromechanical contraptions in your service panel use a combination of methods to monitor the current that passes through them, and if they detect excessive current they open the circuit. The electrical transmission system has similar protections on a larger scale: on the distribution wires strung on poles outside of your house, for example, there are various fuses and circuit-breaker-like devices known as reclosers. High-tension transmission lines 3, like the 230 kV system built by the BPA, need similar protection for similar reasons—except that it is much more complex.

Electrical terminology can be complicated on a good day, and this situation is even more complex because of the historic bifurcation between terms and practices in building electrical wiring versus electrical distribution (which are governed, for example, by two separate electrical codes) and the fact that we are talking about a system that is nearly 100 years old. Relays were a newer technology in the 1930s, as was large-scale over-current protection, so transmission engineers viewed circuit breakers as just an application of the relay and over-current protection on power distribution is still referred to as "relaying" today. Since the whole broader field of supervising transmission lines for safety and reliability is called "protection," relays that open to protect generators, lines, or loads from dangerous conditions are called "protective relays."

Some of the protective relays used in transmission are very much like the circuit breakers in your home. Directional over-current relays, for example, monitor the current passing through them in one direction ("towards" the load) and open if it is excessive. Ground fault relays open when they detect, via various current transformers, that an excessive amount of power is leaking to the ground—just like the GFCI outlets or breakers installed in wet areas of homes.

Some of them, though, are much trickier. The first major problem is directionality. In your home electrical wiring, there is a clear sense of where power flows "from" (the service panel) and "to" (an outlet or fixture). Wiring thus only needs over-current protection in one direction. In a wide-area transmission network, this isn't true. The Master Grid was designed to incorporate a ring for much the same reason that SONET and other communications technologies favored rings: with a ring topology, you can lose the connection between two points and still be able to serve all points by sending power the "other direction." In general, it is common that electrical transmission lines can be "fed" from both ends, and have "load" on both ends. This flexibility to reach the same places by different routes makes the grid more reliable and responsive to changes. It also makes over-current protection more challenging.

Say that you have a span of transmission line, and that somewhere along it a tree falls and pushes one conductor against another. You now have a short-circuit fault at 230 kV (or more in later lines), a dramatic and dangerous condition. You also have thousands, if not hundreds of thousands, of customers that are depending on the power provided by that line. Current transformers can measure the enormous fault current, and indeed it will be detected at many points along the line. But what do you do?

Ideally, protective relays should open on both sides of the fault, and as close to the fault as possible. This cuts power to the dangerous situation while minimizing the number of customers who experience an outage. It's also difficult to achieve in practice: you cannot simply open a protective relay on over-current, or every relay on the line will open at the same time. Instead, analog circuits were used to measure the impedance to the fault as a proxy for its distance. This way protective relays could be carefully tuned to open only when a fault was near them.

Now, consider that a transmission line may be "tapped" and connected to load centers or generation at multiple points along its length. There may also be multiple parallel routes that current can take, with varying capacities. Both of these situations mean that it is often necessary to measure the current (to detect faults) at locations remote from the actual protective relays, which are large devices that needed to be located at substations. Further, the complications of parallel routes and possible ground faults on long lines required the use of a "differential protective relay" or "balanced current relay" that took a much higher level approach to the problem. In a differential system, you measure the current at every connection point to a given protection zone and sum them. The sum should be zero: the same amount of current goes into the line as comes out. If it's not, something has gone wrong somewhere, likely current that is escaping to ground or traveling on a parallel route not engineered for such abuse. But the points where you measure current may be many miles apart, and you still need to make real-time comparisons between them.

In 1940, the BPA was already planning the installation of "pilot relaying" over their carrier current telephone system. In an engineering context, a "pilot" is usually something small that controls something big. A pilot operated relief valve (PORV), for example, is an arrangement where dangerous pressure levels (of a liquid or gas) cause a small pilot valve to open which then triggers a pressure differential in a much bigger valve that causes it to open. There is a similar idea in protective relaying: a pilot-operated relay is a relay that disconnects a very big wire under the control of a smaller wire carrying a pilot signal. The simplest scheme works like this: a device, like a current transformer, monitors a safety parameter and produces a tone whenever it is acceptable. Elsewhere, a protective relay monitors that tone. If the tone ever goes away, the relay opens. The tone might go away because the pilot device detected an unsafe condition, but it might also go away because the line carrying the pilot signal was damaged, making it a fail-safe design. This is good for safety, but bad for reliability, and requires that the communications infrastructure used for protective relaying be very reliable.

Unfortunately, it was not: the carrier current systems installed by BPA through the 1940s were typical of the technology used in the industry at the time, but it was ill-equipped for the scale of the Master Grid. In part to alleviate concern of federal competition wiping out private utilities, and to improve general efficiency, the BPA formed the Northwest Power Pool in 1941. Initially made of about a dozen electrical utilities in the Pacific Northwest, the Power Pool formalized a set of arrangements by which utilities would buy and sell power among themselves, carried over the BPA's transmission network for a small "wheeling" fee. The Northwestern Power Pool would eventually become the Western Power Pool and a template for much of the nation's electrical industry. It significantly increased reliability and efficiency in the region by allowing utilities to sell their overproduction to utilities with high demand, and vice versa. It also brought the carrier communications system to its knees: a practical requirement of the power pool arrangement with wheeling over the BPA's transmission lines was that those transmission lines had to carry protective relaying pilot signals for all of the utilities involved.

A severe fault in a utility taking power off of a BPA line, for example, might require opening protective relays at a power plant operated by a different utility somewhere else on that line. New techniques for communications-aided protective relaying were under development, things like "permissive underreaching transfer trip" and "permissive overreaching transfer trip" that are difficult to explain briefly. These required that protective equipment at each substation have information about the state of protective equipment at all of the other substations, in order to make decisions that are not just based on local conditions but are globally optimal for the health of the whole line.

For example, avoiding a dangerous "islanding" condition on a transmission line might require opening relays at three or more locations along the line, but opening any relays beyond those required would simply cause unnecessary outages. Further, reliability is key and some types of faults on high-voltage lines are self-clearing (this is sort of a euphemism for the fact that a tree branch bridging sufficiently high-tension conductors will often be knocked off of them, if not vaporized entirely, by the resulting arc). Some protective relays are "reclosing" (and are often referred to in brief as "reclosers"), meaning that they will automatically reset (close) after a brief wait period. Some of the time, the fault will be gone and service is restored. Reclosing is potentially dangerous, though, and especially in transmission networks reclosing is only desirable in certain circumstances. Further pilot signals can be used to enable and disable reclosers based on the nature of the fault or the other locations at which it was detected, so that for example a fault that has taken a power plant offline does not lead to a recloser elsewhere "flapping" and stressing the remaining production capacity.

By the end of the 1940s, the BPA's carrier telephone system was overstressed to the point of failure. Power Pool utilities had connected their own carrier equipment, butting frequency bands so closely against each other that they began to interfere and degrade the quality of connections that could already be difficult to make out over hundreds of miles of high voltage infrastructure. The fact that a fault on a transmission line would generally prevent carrier current communications over that line working was a feature for simple fail-safe pilot relaying systems, but as protective relaying technology advanced it became the source of cascading failures. For example, by the early 1950s the BPA had found that lightning strikes on major transmission lines would cause enough interference to the carrier current system that protective relays all along the line would lose their pilot signals, open as fail-safe, and escalate what should have been a localized problem into a systemic one. Besides, many sections of the network had by that time become so congested by various carrier current circuits that there was simply no room in the spectrum for more—and so no capability to install protective equipment for new service lines.

Microwave

Meanwhile, the Second World War had had a profound impact in two ways: first, wartime demand for weapons, aircraft, and aluminum had driven Pacific Northwest industry to new heights. In 1949, power consumption in the Pacific Northwest had more than tripled, and population increased by 45%, compared to 1939. The Corps of Engineers built more dams, so the BPA's network carried more power. Engineers at the Administration's laboratory developed the techniques to operate transmission lines at record-setting voltages, 300 kV and beyond—making protection systems all the more critical. Second, wartime radar research had produced technological byproducts including high-bandwidth microwave transmitters and receivers.

So, the electrical industry adopted microwave communications technology alongside the telephone industry, and for much the same reasons. The high bandwidth of microwave links allowed for multiplexing huge numbers of channels, and the lack of wireline infrastructure (especially shared with the actual transmission lines) promised increased reliability. In my previous article on passive repeaters, I noted that they were especially popular with electrical utilities. Now, we learn why: utilities throughout the country built microwave communication systems that carried a fraction of the traffic of the Bell System, but often carried it to more remote locations and with higher reliability requirements.

The BPA's situation was different from that of more traditional utilities. Most private electrical utilities had started in a city and grown outwards, with a denser service network and fewer long-distance lines. They had mostly built private networks for communications, stringing their own open-wire telephone lines between stations. The BPA, with its far-flung network and huge distances, couldn't afford that kind of investment in stringing wires. In some ways, this proved an advantage: they could start from a clean slate. This network would be architected from the beginning for efficiency and performance, rather than to accommodate existing infrastructure.

While the BPA had initially built substations for a larger staff, the control technology was rapidly evolving and some of the more intensive activity BPA had expected at substations proved unnecessary 4. Despite the generously sized control houses, by 1950 most substations had only a single operator on staff, who would often be "out in the field" tending to the transmission lines. This created a problem if some sort of sudden problem required reconfiguring the network to restore service—a system operator might be left ringing a substation phone with no one there to answer.

That year, three substations had been equipped with "supervisory control" technology. This new system combined telemetering and teleoperation, so that a system operator in Portland could monitor voltage and current measurements at the substation and remotely operate the switchgear. The benefits of supervisory control were obvious, but the limitations of the carrier telephone system meant that those three substations were as far as the system could reach. Based on its promising experience with supervisory control of those three stations, and the clear need to continue expanding an already-stressed communications system, the BPA in 1949 committed to the construction of a completely new, completely modern control system for the Northwest Power Pool.


I put a lot of time into writing this, and I hope that you enjoy reading it. If you can spare a few dollars, consider supporting me on ko-fi. You'll receive an occasional extra, subscribers-only post, and defray the costs of providing artisanal, hand-built world wide web directly from Albuquerque, New Mexico.


The BPA's substation in Vancouver, already one of the largest, had a combination of ample space and proximity to Portland that made it a convenient location for support facilities. Construction yards, maintenance shops, and research laboratories had all been added onto the facility. The 1949 project launched with a symbolic gesture: the site was renamed, from simply the North Vancouver Substation to the J. D. Ross Complex in honor of the BPA's first administrator. A new building at the Ross Complex, the Control Center, became the nerve center of the system and the Rome to which all microwave routes led.

Keeping with our communications theme, one of the main features of what was then called the Ross Control Center was a "three position turret" (yes, a turret!) by which operators could call any line on the microwave, carrier current, or leased line telephone networks. The plan was to obsolete the turret, though. In 1950, the BPA awarded a half-million-dollar contract to the Philco corporation (originally Philadelphia Battery and later part of Ford, then GTE, then Philips, amusingly providing another expansion of the name) for its new microwave network. Philco had done extensive military work on microwave radar during the war, and was at the time one of the leaders in microwave technology. Philco's expertise would be needed, because the microwave network ordered by the BPA would be the largest electrical utility communications network in the world. While more difficult to conclusively state, I think it is likely that it was either the second largest microwave network of any kind after AT&T's, or the third largest after those of AT&T and the Santa Fe Railroad (both of which were also Philco customers).

At this point, the BPA's transmission network had extended to the Hungry Horse Dam in northwestern Montana, adding customers along the way. There were some 14,000 circuit miles of transmission lines and 400 substations in the system, and Philco's contract specified that they would build out the initial operating network in just 360 days. Work got underway: BPA signed the contract in February of 1950, in September testing and demonstrations were conducted, and in October the BPA officially activated the first leg: a 200 mile route from the Ross Complex to Snohomish, Washington.

Snohomish

This first route, built at a cost just under a million dollars, used repeaters at Mt. Rainier, Chehalis, Olympia, Squak Mountain. The longest single jump was Olympia to Squak Mountain, about 55 miles, an unusually long distance for microwave facilitated by Squak Mountain's prominence and a 150' tower at Olympia. With multiplexing equipment, this route carried 23 channels from the control center to the Puget Sound area.

This first microwave link was quickly put to work for one of the most interesting new applications of utility telecommunications: fault locating. When a fault occurred on a long transmission line, the BPA's first step was to search the whole line for the problem. Many lines were in difficult terrain, so a helicopter or airplane was used to speed up the process. The faults were sometimes minor and not easy to see from the air (say, a broken insulator), so the survey aircraft would take photos for processing and analysis on the ground. This process was expensive and, moreover, it was time consuming. BPA engineers realized that sudden open circuits or shorts in transmission lines created electrical signals that propagated back through the line and could be observed on test equipment—so you could presumably locate a fault by calculating its time of flight to the substations at each end. The problem was obtaining a measurement of when the fault signal arrived at two different locations, in precise synchronization.

Requiring high speed and, more importantly, consistent latency, this was exactly the kind of problem that microwave lines were well suited for. ITT 5 designed the system to BPA specifications, including devices that detected fault waves and reported them over a microwave channel, and a machine that compared the timing of the received reports and calculated a likely fault position (in miles) relative to each substation. FTL seems to have estimated that the system was accurate to 600', BPA to 1,000'.

The Ross-Snohomish route carried other important traffic as well: as part of its inaugural celebration, Washington State Representative Henry M. Jackson used the new internal telephone at Snohomish to call Ross over the microwave link, congratulating BPA's administrator and chief engineer on the accomplishment. Although I am unclear on the exact criteria being used, newspaper reports consistently identify the link as the "first of its kind in the world" 6.

1950 was still early for microwave technology; AT&T's first commercial microwave link had only gone into service in 1948, and that was experimental. The transcontinental telephone "skyway" wouldn't be completed until a year later. As a result, microwave technology was unfamiliar to the public, and the appearance of parabolic antennas on BPA facilities—and at repeater stations on mountaintops and out in the woods—was conspicuous. One reporter called them the BPA's flying saucers, another explained the repeaters in terms of "pitcher" and "catcher." Every paper ran photos.

Over the next two years, Philco completed a second microwave route up the Columbia Gorge connecting each of the dams through to Spokane, and a third that linked Beverly, on the route to Spokane, to Snohomish—forming a ring like the original Master Grid that provided redundancy and direct protection channels for transmission lines on that route. The 1952 microwave network included fourteen primary terminals and 21 repeaters; it connected the dispatch telephone system at Ross with dams and substations along the microwave routes as well as seven mountaintop HF stations to reach field crews.

As was typical at the time, the microwave equipment at each repeater and terminal ran directly from battery power. The batteries were charged (normally floated) from two different power supplies, one from the utility and the other from an on-site propane generator with a two week fuel supply. BPA initially used prefabricated aluminum shelters for the equipment at repeaters, although many were originally built or later rebuilt as cinderblock. This was, in part, due to the weather: mountaintop repeater stations in Washington coped with severe winters, and BPA went through several rounds of modifications to their building and tower designs. Towers were reinforced against ice accumulation, and repeater stations in particularly snow-prone parts of the Gorge and the Snoqualmie Pass were made two story. These buildings had a "balcony" entrance on the second floor with a ladder to reach it, allowing access even when the first floor was completely buried in snow. Towers were rated for 100 MPH winds, and some for thousands of pounds of ice. BPA designed ice shields for the antennas, and a heated cover to keep ice from covering the reflector surface.

BPA also took advantage of passive repeaters to relocate microwave sites to more accessible locations. At Rockdale, on the edge of the Cascade Mountains, a repeater was installed just next to the highway. Its antennas aimed more upwards than sideways, at reflectors at the top of the mountain ridge. Microflect passive repeaters were manufactured in Oregon, and the BPA was one of Microflect's first large customers, contributing design changes for mountainous service.

By 1955, the BPA microwave network had reached Hungry Horse Dam in Montana and connected all of the major substations of the system. When a decision was made, in 1955, to relocate the power dispatch office from the Ross Complex to the new Portland headquarters, the capacity and expandability of the microwave system made the process much easier. The BPA's new HQ building looked more like a telephone exchange than a federal building: one of its most prominent features was a rooftop microwave tower. The 1957 annual report inventoried 61 microwave radio sites covering 1,300 miles of route, by then using a mix of Philco, ITT, and Motorola equipment.

BPA's network had, by this time, achieved many feats of rural service. One of the most impressive was service to the southern Oregon Coast: this area was very remote and faced terrible weather throughout the winter. BPA's 115 kV South Coast line was the only electrical service into the region, and it was regularly disabled by ice storms and flooding. Because of the area's mountainous terrain, line crews working in the area were only infrequently able to reach a substation near Eugene on their mobile radios, and that was their only way of talking to dispatchers to coordinate repairs. In an effort to improve service reliability, the South Coast became the pilot for a new model of field communications. Relatively closed spaced microwave repeaters, each with a VHF radio, would bridge radio channels onto the microwave telephone system. To stand up to the Oregon Coast, each of these stations used an aluminum enclosure with heating, fire suppression, and a generator. Some of the enclosures were cabled to the ground, to better hold them down against the wind.

During the 1960s, the value of the microwave system had been proven but it was once again facing the limits of its capacity. Microwave technology had improved tremendously in the post-war decade and BPA's 23 and 24-channel multiplexers were obsolete. A $1.6 million contract was let to Lenkurt to upgrade much of the microwave network to the modern Lenkurt 76C microwave radio and 46A or 34A multiplexers, capable of up to 600 channels at around 8 GHz (the previous system varied from site to site but operated at around 2 GHz).

This was equipment designed and built under GTE ownership, and was also typical of the long-distance microwave links in the GTE telephone network. As part of the project, Lenkurt also installed VHF radio relays throughout the system. At some sites, Lenkurt installed dual-polarized antennas to allow simultaneous operation of the old and new multiplex systems and, later, increased capacity. Further contracts expanded the microwave network to Bellingham, Washington and to Corps of Engineers projects on the Snake River. In 1966, BPA dispatchers in Portland could monitor production at 21 dams, receive alarms from 250 substations, and completely remote control fifteen of the network's key switchyards.

When the BPA built the Pacific Intertie, a combination of two 500 kv AC circuits and one 800 kV DC circuit stretching 900 miles from the Columbia River to near Los Angeles, it was the largest transmission line project in US history. Under construction from 1965 to 1970, the line's standard bearer came in microwave form. Collins Radio built microwave routes parallel to the Intertie transmission lines, a $2 million project with 22 new radio stations on the 600-channel Lenkurt system. Most of the work was finished by 1967, a prerequisite for some construction the transmission line itself, since the microwave channels were used for testing and commissioning.

The Computer Age

Microwave was not the only arena in which the mid-century had brought new technology. The BPA was not new to computers; various forms of electromechanical computation had been part of their engineering works since the 1930s. By the 1960s, though, projects like SAGE (a military air defense system) and SABRE (a commercial airline reservation system) demonstrated the potential of combining computers with telecommunications for real-time control. The BPA had a telecommunications network, and it had a real-time control system... and they decided to add a computer.

In 1966, the BPA announced that its control center would once again move, from the Portland headquarters building back to the Ross Complex, where a new building would be designed from the ground up for centralized, computerized control of the power system. Named the Dittmer Control Center after a previous BPA power manager, the low-slung building had a prominent concrete microwave tower and was partially sunk below ground level for hardening against attack (it was, after all, the Cold War).

Much of the Dittmer center's lower-level floorspace was devoted to equipment rooms, which would soon be the home of the computer system that BPA contracted to Rockwell. Among other equipment, Rockwell installed a PDP-10 computer that received, analyzed, and logged telemetering data from throughout the system for display to operators. From its commissioning in the early 1970s, BPA continued to enhance the computer center into an integrated power dispatching system that supported operators in monitoring the network, predicting future demand, switching transmission lines, and ordering changes in production at power plants throughout the Pacific Northwest.

The ultimate manifestation of the PDP-10's software was called RODS, the Real-Time Operations, Dispatch, and Scheduling System. RODS was one of the first systems of its kind, initially contracted to Rockwell in part due to their experience with control computers for the Apollo program. Features of RODS included a time-synchronized data acquisition system to support differential current monitoring (ensuring that the computer compared current measurements taken at precisely the same time), which used an atomic time standard at the Dittmer Control Center to distribute high-precision timecode through the microwave network. As RODS matured, it established the 15-minute scheduling loop used by dispatchers to configure power plants and transmission lines for a constantly changing electrical load. DEC themselves, in an internal sales meeting whose minutes fortunately made it into the historic record, noted the Dittmer PDP-10 as a critical early sale in their efforts to break into the utility industry.

Despite the technical firsts of the Dittmer Control Center and RODS, this chapter of BPA history is best known for its incidental brush with the Pacific Northwest's most famous chapter of computer history: as temporary employees of TRW, the aerospace contractor brought on for RODS software development, Bill Gates and Paul Allen both spent time at Dittmer. They were still in high school, it would be years before they moved to Albuquerque to found Microsoft.

Many parts of the BPA's microwave network are still in use, although improving radio technology and the adoption of fiber optics (including fiber embedded into the neutral conductors of transmission lines) have allowed for elimination of some repeaters. As microwave technology continues to fall obsolete (in comparison to fiber, commercial terrestrial radio networks, and satellite), many of the remaining sites are likely to be demolished. Some of them—sites like Chehalis, Squak Mountain, and Rainier—have been in service for 76 years.

The BPA microwave network is not unusual. It was the first of its type, but the transmission and power marketing concepts pioneered by the BPA are now used nationwide. Wherever power goes, protective relaying follows, and the telecommunications networks that shadow the electrical grid from coast to coast. Few enterprises outside of the communications industry itself have ever operated communications networks on the scale of electrical utilities. They find themselves in the company of railroads and oil pipelines, ventures that must span deserts, climb mountains, and cross rivers.

Unlike many of those operations, though, the BPA is a federal agency, subject to the National Environmental Policy Act and National Historic Preservation Act. BPA's extensive Section 106 NRHP compliance program has produced extensive historic documentation of its transmission lines, communications infrastructure, and research facilities. While sometimes onerous, these federal policies have created the best-documented historic electrical utility communications system in the nation.

An inventory of historic microwave stations still under BPA ownership turned up 28 in Washington, 22 in Oregon, one in Idaho, and two in Montana, and many are eligible for nomination to the Historic Register. From an aluminum shed at the Ross Complex to a squat concrete building and utility pole at Mary's Peak, they are monuments to American history—a singular, almost megalomaniacal vision of the Columbia River tamed; the rise of Pacific Northwest industry under wartime demands; a technical approach to environmental preservation and economic equality. Infrastructure for the public benefit: a great American achievement. We might, one day, find the will to do it again.

  1. This Vancouver is just across the Columbia from Portland, and should not be confused with the big one in British Columbia. Having grown up in Portland, I will probably not disciplined enough to put Washington after it each time, so just remember that we aren't talking about Canada. Yet.

  2. Lenkurt would later merge with GTE, becoming something like the Western Electric to AT&T's principal competitor.

  3. This is another example of the difficulty of electrical terminology. "High tension" in this context is a result of the historic use of "tension" (and enduring use in some languages) to refer to what we now usually call "voltage" or "potential." Specifically in electrical distribution, "high tension" is older language but still in use to refer to 100 kV and higher transmission lines.

  4. BPA was navigating many firsts in the construction of the Master Grid, and the high-voltage transformers used to step up and down from the 230 kV lines were new technology. They were filled with a heavy oil for cooling and electrical insulation, and the BPA had at first expected that they would need to drain the oil and open the transformers on a regular basis. Early BPA substations featured a network of rails and low-slung metal carts that would be used to pick the transformers up off of their pads and move them to a specialized building called an "untanking tower," where they could be drained of oil and the covers lifted off by a gantry crane. In practice, transformers were virtually never opened, and the untanking towers and rails became vestigial.

  5. This company changed names from Federal Telecommunication Laboratories to Federal Telephone and Radio and was then acquired by International Telephone and Telegraph (ITT), all during the time period covered by this article. I will refer to it simply as ITT for readability.

  6. You know that I am a little bit obsessive about this kind of superlative, and since it seems to have originated with Chief Engineer Sol Schultz, I tend to think he had a good definition in mind. My best guess is that it was the first microwave link in the world (at least that they were aware of) to carry telemetering and teleoperation signals over such a long distance.

Error'd: The Song that Never Ends?

"Watch to the end", the scammers demand. In today's episode of Error'd, the end is a long time coming. But at least it's amusing. In the meantime, Amazon irked and/or terrified thousands of their customers last week by mailing out ridiculously inflated bills. It's been covered extensively elsewhere but why should we miss all the fun?

It was Willy who worried "Amazon prices seem to have crept up this month, finance are going to have something to say"

a7e89b1dc59e4c3ba3b8571a056825d6

Dave A. is on the horns of a dilemma. "Is it optional or required? Make up your mind, TAP! I'm TAPping my fingers waiting for you to decide."

ffd147510c4b455c9c43e60763e68bf8

"This rating is through the roof!" exclaims a regular who wants to be Anonymous today. "We're done with linear rating scales. Now you can have 5 stars on both X and Y axes. Also, given this is for a company that cleans roofs, we can make a joke this rating is out of the roof, and out of the <div> as well. In case you're interested and in case you want take a screenshot yourself: https://mijndakschoon.nl/" It's real; I checked.

11083f7140b449a48c5218843699cc24

Richard H. found a flubstituted email. "Quickbooks here reminding me that I failed to pay invoice number "{{rand_invoice}" Or this might be a scam/phishing email. Hard to tell." I'll bet {{money}} on scam.

710be181fec049ab8f717ae185783f5f

B.J. H. is usually quite concise. Usually. "If this is ALL NORMAL I'm worried what will happen in an emergency"

7885ce2a1ae64033b9b46e472925edd9

[Advertisement] Plan Your .NET 9 Migration with Confidence
Your journey to .NET 9 is more than just one decision.Avoid migration migraines with the advice in this free guide. Download Free Guide Now!

Classic WTF: My Many Girlfriends

Honestly, with the wildfire smoke and the oppressive heat, maybe it's time to find a nice quiet cave to hang out in. Something with no natural light and no natural ventilation. I wonder if anybody has a place like that… Original. --Remy

In the long ago, wild-west days of the late 90s, there was an expectation that managers would put up with a certain degree of eccentricity from their software developers. The IT and software boom was still new, people didn't quite know what worked and what didn't, the "nerds had conquered the Earth" and managers just had to roll with this reality. So when Barry D gave the okay to hire Sten, who came with glowing recommendations from his previous employers, Barry and his team were ready to deal with eccentricities.

Of course, on the first day, building services came to Barry with some concerns about Sten's requests for his workspace. No natural light. No ventilation ducts that couldn't be closed. And then the co-workers who had interacted with Sten expressed their concerns.

During the hiring process, Sten had come off as a bit odd, but this seemed unusual. So Barry descended the stairs into the basement, to find Sten's office, hidden between a janitorial closet and the breaker box for the building. Barry knocked on the door.

"Sten awaits you. Enter."

Barry entered, and found Sten precariously perched on an office chair, removing several of the fluorescent bulbs from the ceiling fixture. The already dark space was downright cave-like with Sten's twilight lighting arrangement. "He welcomes you," Sten said.

"Uh, yeah, hi. I'm Barry, I'm working on the Netware 3.x portion of the product, and Carl just wanted me to check in. Everything okay?

"This is acceptable to Sten," Sten said, gesturing at the dim office as he descended from the chair. Sten's watched beeped on the hour, and Sten carefully placed the fluorescent bulb off to the side, in a stack of similarly removed bulbs, and then went to his desk. In rapid succession, he popped open a few pill containers- 5000mg of vitamin C, a handful of herbal and homeopathic pills- and gulped them down. He then washed the pills down with a tea that smelled like a mixture of kombucha and a dead raccoon buried in a dumpster.

"He is pleased to meet you," Sten said, with a friendly nod. Barry blinked, trying to track the conversation. "And he is pleased with it, and has made great progress on building it. You will like his things, yes?"

"Uh… yes?"

"He is pleased, and I hope you can go to him and tell him that he is pleased with this, and set his mind at ease about Sten."

So it went with Sten. He strictly referred to himself in the third person. He frequently spoke in sentences with nothing but pronouns, and frequently reused the same pronoun to refer to different people. The vagueness was confounding, but Sten's skill was in Netware 2.x- a rare and difficult set of skills to find. So long as the code was clear, everything would be fine.

Everything was not fine. While Sten's code didn't have the empty vagueness of unclear pronouns, it also didn't have the clarity of meaningful variable names. Every variable and every method name was given a female first name. "Each of these is named for one of Sten's girlfriends." Given the number of names required, it was improbable that these were real girlfriends, but Sten gave no hint about this being fiction.

There was some consistency about the names. Instead of i, j, and k loop variables, you had Ingrid, Jane, and Katy. Zaria seemed to be only used as a parameter to methods. Karla seemed to be a temporary variable to hold intermediate results. None of these conventions were documented, obviously, and getting Sten to explain them was an exercise in confusion.

It led to some entertaining code reviews. "Michelle here talks to Nancy about Francine, and then Ingrid goes through Francine's purse to find Stacy." This described a method (Michelle) which called another method (Nancy), passing an array (Francine). Nancy iterates across the array (using Ingrid), to find a specific entry in the array (Stacy).

Sten lasted a few weeks at the job. It wasn't a very successful period of time for anyone. Peculiarities aside, the final straw wasn't the odd personal habits or the strange coding conventions- Sten just couldn't produce working code quickly enough to keep up with the rest of the team. Sten had to be let go.

A few weeks later, Barry got a call from a hiring manager at Initrode. Sten had applied, and they were checking the reference. "Yes, Sten worked here," Barry confirmed. After a moment's thought, he added, "I suggest that you bring him in for a second interview, and have him walk you through some code that he's written."

A few weeks after that, Barry got a gift basket from the manager at Initrode.

Thanks for the tip

Sten did not get hired at Initrode.

[Advertisement] Picking up NuGet is easy. Getting good at it takes time. Download our guide to learn the best practice of NuGet for the Enterprise.

Classic WTF: Server Room Fans and More Fun

It's been pretty hot lately. Probably should use a fan to cool off. Mind the trip hazards. Original. --Remy

"It's that time of year again," Robert Rossegger wrote, "you know, when the underpowered air conditioner just can't cope with the non-winter weather? Fortunately, we have a solution for that... and all we need to do is just keep an extra eye on people walking near the (completely ajar) server room door."

 

"For as long as anyone can remember," Mike E wrote, "the fax machine in one particular office was a bit spotty whenever it was wet out. After having the telco test the lines from the DMARC to the office, I replaced the hardware, looked for water leaks all along the run, and found precisely nothing. The telco disavowed all responsibility, so the best solution I could offer was to tell the users affected by this to look out the window and, if raining, go to another fax machine."

"One day, we had the telco out adding a T1 and they had the cap off of the vault where our cables come in to the building. Being curious by nature, I wandered over when nobody was around and wound up taking this picture. After emailing same to the district manager of the telco, suddenly we had the truck out for an extra day (accompanied by one very sullen technician) and the fax machine worked perfectly from then on."

 

"I found this when I came back in to work after some time off," writes Sam Nicholson, "that drive is actually earmarked for 'off-site backup'. Also, this is what passes for a server rack at this particular software company. Yes, it's made of wood."

 

"Some people use 'proper electrical wiring'," writes Mike, "others use 'extension cords'. We, on the other hand, apparently do this."

 

"I was staying at a hotel in Manhattan and somehow took a wrong turn and wound up in the stairwell," wrote Dan, "not only is all their equipment in a public place (without even a door), it's mostly hanging from cables in several places."

 

"I spotted this in China," writes Matt, "This poor switch was bolted to a column in the middle of some metal shop about 4m above ground. There were many more curious things, but I decided to keep a low profile and stop taking pictures."

 

[Advertisement] Keep the plebs out of prod. Restrict NuGet feed privileges with ProGet. Learn more.

CodeSOD: Classic WTF: Fork and Log

We keep our summer break going. Today, there's something floating in the pool, and I don't think it's a Snickers bar. Original. --Remy

A few years back, Adam C. was brought in to help with some performance problems that appeared while load testing a VXML Platform. The project was already well behind and they couldn't figure out why the system kept falling over under a very slight load. To make matters worse, Adam had absolutely no prior knowledge of the system or its software other than Wikipedia’s definition of what VXML is.

A veteran to these sorts of situations, Adam grabbed a coffee, a donut, and then started picking through the application logs to get a feel for what the system is doing and where something might be going wrong.

Looking in /var/log/messages, he was pleased to find with several days’ worth of messages, but over and over again, the same entry popped up:

Exception encountered writing error log. 

"Seriously, who logs an error that they can't log an error? ...and is that even possible?" Adam wondered aloud after seeing the same senseless message for what he figured to be the hundredth time.

Frustrated, and hoping to learn what ludicrous conditions might precipitate a log to contain such a message, Adam dug into the source and hit paydirt in the form of this wonderful nugget of code:

public void error(String logID, String errStr) {
  StringBuffer errLogCmd = new StringBuffer("/usr/bin/logger -p ");
  try {
    Runtime rt = Runtime.getRuntime();
    errLogCmd.append(errlogFacility);
    errLogCmd.append(" -t ");
    errLogCmd.append(logID);
    errLogCmd.append(" ");
    errLogCmd.append(errStr);
    rt.exec(errLogCmd.toString());
  } catch (Exception ele) {
    System.out.println("Exception encountered writing error log." + ele.getMessage());
  }
}
As he mentally parsed his way through the code, Adam could feel his breakfast tickling up from the back of his throat in reaction to the number of WTF’s he found himself facing.

He couldn’t decide what about the implementation was worse - forking off an external process to log something, the fact that log4j could have been used to send stuff to syslog (which the project used elsewhere), or that since the program's output was already piped to /usr/bin/log, just doing a System.out.println() would have been equivalent to this code.

As it turned out, this wasn’t the root cause behind the performance problems, but needless to say, that code got junked rather quickly and he moved on to looking for the next performance bottleneck.

[Advertisement] Plan Your .NET 9 Migration with Confidence
Your journey to .NET 9 is more than just one decision.Avoid migration migraines with the advice in this free guide. Download Free Guide Now!

CodeSOD: Classic WTF: The Table Selector

It's summer break time, which as always, means we dip back into classic articles. Today, we pick which table we want. Original. --Remy

"In my native language of German," writes Christian, "the word quellcode is a pretty direct translation of 'source code'."

"Unfortunately, bad code seems to cross language barriers - as does that famous three-letter explicit adjective. But occasionally I’ll find a piece of quellcode that deserves its own special, localized expletive: quäl-kot. When I stumbled across this interface in our quellcode, quäl-kot was the first thing that came to my mind."

public interface ITableSelector
{
    string selectTable1();

    string selectTable2();

    string selectTable3();

    string selectTable4a();

    string selectTable4b();

    string selectTable5();

    string selectTable6();

    string selectTable7a();

    string selectTable7b();

    string selectTable8();

    string selectTable9a();

    string selectTable9b();

    string selectTable10();

    string selectTable11();

    string selectTable12();

    string selectTable13();

    string selectTable14a();

    string selectTable14b();

    string selectTable14c();

    string selectTable14d();

    string selectTable15();

    string selectTable16();

    string selectTable17();

    string selectTable18();

    string selectTable19();

    string selectTable20();

    string selectTable21a();

    string selectTable21b();

    string selectTable22();

    string selectTable23();

    string selectTable24();

    string selectTable25();

    string selectTable26();

    string selectTable27();

    string selectTable28();

    string selectTable29();

    string selectTable30();

    string selectTable31();
}
[Advertisement] Picking up NuGet is easy. Getting good at it takes time. Download our guide to learn the best practice of NuGet for the Enterprise.

Error'd: Princess Pricing

Sam suggests this Error'd indicates "Disney+ preparing the ground for usage-based billing." I'm intrigued by the idea that Disney might charge by the minute, but I suspect the reality is far more mundane.

13e08caf609f4133afd419a76acb432a

Silly prices at online shopping sites don't usually make it through the gauntlet here, but I'm making an exception for the math, as Rob H. points out "it ain't mathin'."

0924416949684290b36ad0d823129181

and Harrison suggests a novel kind of discounting math "I went to the supermarket later at night for some beers, and had a snoop around the yellow sticker items for anything I might need or could freeze. This bakery item was priced in reverse, Was: 0.00, Now 1.99, discounted by negative infinity percent, and infinity is even printed upside down somehow."

d41527fab58d46fa97f4477dfcbf659d

We've got a mojibake from dragoncoder047: "Was browsing through the widget options on my iPhone home screen and found that Game Center had decided to do this. Mind you, my iPhone was, and always had been, set to English."

c4fb48f719b34b5393e274cc6bc48f2a

Finally a combination of typical time travel and package tracker shenanigans, not explained by time zone hijinks. Evelyn notes "Apparently the package was registered in July, on its way during January, and then got back to July."

ef7d6784e50a48eebfaf4b912c8caa55

[Advertisement] BuildMaster allows you to create a self-service release management platform that allows different teams to manage their applications. Explore how!

CodeSOD: Wait Longer

Karen was maintaining some specification tests that were flaky. Not extremely flaky, but three or four times out of a thousand, the tests would just fail. The tests were complicated, and some of the operations were timing sensitive, so it wasn't precisely surprising- but the problem was that they were actually generous with their timing windows. The unit tests passed consistently, it was only these functional, specification-based tests that failed.

So, for example, there were sections in the tests where they wanted to wait at least 2ms. Since the code and tests were in TypeScript, they used the setTimeout function, which per standard JavaScript documentation warns that it may wait longer. But again, Karen was fine with longer.

Unfortunately for Karen, the documentation for NodeJS is less specific, as it makes no guarantees about when the timeout function gets invoked. This means that it can fire the timeout before the time has elapsed.

After many, many hours of debugging, that was exactly the situation that Karen found herself in. Which is why her very simple wait function went from:

export const wait = (ms:number) => new Promise((complete) => setTimeout(complete, ms));

To the much more awkward:

export const wait = (ms: number) {
    const target = performance.now() + ms;
    return new Promise((complete) => {
        const checkReady = () => {
            if (performance.now() > target) {
                complete();
            } else {
                setTimeout(checkReady, 1);
            }
        }
        setTimeout(checkReady, 1);
    });
}

This version of the function checks the time every millisecond, and only completes the operation if we've waited at least as long as our target duration. This ensures that the timeout never fires too soon and it fixes the janky tests. But it's also terrible. Terrible that it exists. Terrible that this is the best solution. Terrible that our functional tests need to be so time sensitive. And terrible that the Node runtime actually breaks the one consistent scheduling guarantee that pretty much every other scheduler does: that it'll wait at least as long as you asked, but might wait much longer.

At best, we can say, "at least it's only testing code."

[Advertisement] Picking up NuGet is easy. Getting good at it takes time. Download our guide to learn the best practice of NuGet for the Enterprise.

CodeSOD: The Error Check

Today's submission is less a WTF and more a, "Yeah, that'd annoy me too."

Stevie works in a code-base that's largely C, which means function return values are usually used to communicate to status codes. The standard:

BOOL success = someFunc();
if (!success) {// handle the error

If someFunc returns TRUE, we succeeded, otherwise we failed.

There's nothing wrong with that convention. But there is something wrong with one of the long-time developers on the project, because they have their own idiom for doing this. And they've been around long enough that their approach is the convention other developers follow. It's not wrong, per se, just confusing:

BOOL error = someFunc();
if (error == FALSE)
{
    //handle the error
    errorCode = GetLastError();
    //…
}

Yes, pretty much anywhere an error can happen, they check if error == FALSE, and if that's true, they have an error.

I'd say, "at least they're consistent", but they're not. It's the convention, sure, but nobody wrote this down as the convention. New developers come in all the time, and they start out writing code in a more "normal" pattern. But the existing code has its pattern, and there is a lot of it. It has a mass and an inertia that is stronger than any developer. They don't change the code, the code changes them.

[Advertisement] BuildMaster allows you to create a self-service release management platform that allows different teams to manage their applications. Explore how!

CodeSOD: AAYFN

Jason M sends us some Ruby code.

def bv(prop, tv="Yes", fv="No", nv="Not Specified")
  v = self.send(prop)
  if v === true
    tv
  elsif v === false
    fv
  else
    nv
  end
 end

The obvious WTF here is the function name and parameter names. AAYFN(always abbreviate your function names) seems to be the convention here. But it also contains a more Ruby-specific WTF.

The function name bv is short for "boolean value", obviously. tv is the "true value", fv is the false value, and nv is the null value. So this is really about pretty-printing boolean values and converting them to strings. While the terrible names make it hard to understand, it's not that hard to figure out what's going wrong here. I hate it, don't get me wrong, but it just makes me sigh with disappointment, not groan.

No, the thing that makes me groan is v = self.send(prop). This is a very Ruby idiom that lets you access a member of the class by name; essentially it's like doing self.prop, but since prop is a variable containing a property name, we have to send it.

This is metaprogramming by strings, which is the main reason I end up hating it. But in this specific instance, it offers us a lot of potential issues. First, if prop is anything not boolean, we just return "Not Specified", which is incredibly misleading. I'd argue that if we attempt to use it on a non-boolean field we should throw an exception. Which opens the question: if prop doesn't exist, is that a non-boolean field, or is that a "enh, just call it null" situation? Because right now, that will throw an exception. It'll also throw an exception if the thing being accessed is a function that takes parameters. These behaviors may be surprising.

Now, I don't know the calling pattern. It's possible that whatever function calls this already has a good list of the allowed boolean values, and will never call this on a prop that doesn't exist. Certainly, that's what the Ruby docs recommend. But I'm going to hate it anyway, because this kind of runtime metaprogramming by passing strings around is eternally asking for trouble. And while I haven't done a huge amount of Ruby, I've done enough to know that any non-trivial codebase ends up like this once the metaprogramming band-aid is taken off.

Now, if you don't mind, I'll go back to doing my metaprogramming with C++ templates, which are simple, clear, and never result in wildly unmaintainable code because you ended up reimplementing LISP in template operations.

[Advertisement] Utilize BuildMaster to release your software with confidence, at the pace your business demands. Download today!

The Easy No

There were thousand tickets in the backlog, and I was on the trail of a weird printer issue. I had a suspect, but that wasn't enough to close the ticket. I'm Anonymous. This is my story.

Whenever you hit the Print button, a request launches into the ether. “Print one copy of this file, double-sided.”

But you can’t just chuck it out there aimlessly. You gotta tell it where to go. So that’s why we have Internet Protocol, or IP. Every computer, every printer, every device on every network has one or more IP addresses for different operations. They’re just like the address you scribble on the envelope holding Grandma’s Christmas card. Send your print request to the right IP, and you’re in business. Send it to the wrong IP, and they’ll be shaking their heads on the other end. “Print? But all I know how to do is tell you the weather!”

As for what happens next, you’re at the mercy of the program’s error handling. If you’re lucky, they’ll send a detailed error message. If you’re not, they could ignore you, act up in weird ways, or completely self-destruct in a violent crash.

I was neck-deep in a Tech Support ticket concerning a printer in Human Resources. Not only was Grandma’s Christmas card landing in Pluto, the Plutonians were writing back, in the form of mysterious complaints that were nothing like the documents the printer had been tasked with.

My troubleshooting hadn’t shot much trouble at all, so I’d bugged some friends for ideas. Reynaldo had been reminded of an old, retired HR program for logging anonymous employee complaints. While our developer friend Megan quit our icy outdoor break spot to dig up dirt from the software side of things, Reynaldo and I went to his cubicle to trace IPs. He suspected a conflict … and that’s just what his command-line requests revealed. We leaned over his laptop together, studying a couple of terminal windows and their cryptic output to his even-more-cryptic inputs.

“Whenever that printer got set up in HR, it received an IP that was already in use by this server—” he jabbed an index finger at the relevant part of the screen “—which is apparently still up and running, even though I swear it was decommissioned.” Reynaldo frowned. “Who wants to bet that stupid complaint program has some kind of ability to print cached complaints?”

I frowned, too. “We won’t know for sure unless Megan digs something up.”

Reynaldo flashed me a look of pure cynicism. It wasn’t an indictment of Megan’s skill, more skepticism toward the idea of anyone having properly documented this mess. We had better odds of hitting a billion-dollar jackpot.

I sighed. “It’s an easy workaround, at least: assign the printer a new IP. But we’ve got bigger problems here. Why’s this zombie server still up and kicking? Why’d an IP conflict ever occur in the first place?”

“Don’t get me started on our FUBAR IP allocation and tracking.” Reynaldo’s lowered voice contained plenty of venom. “We’re talking multiple pools to track, multiple spreadsheets that have to be manually edited. Problems that are easy enough to fix! Why spend money on honest-to-goodness network management software?”

Short-term budgeting preventing long-term improvements: a damned shame we’d both seen all too many times. Being a tech support drone, having an innate desire to help that no amount of baloney could stamp out, I once more found myself longing to fix the unfixable.

“In my ticket notes, I’ll make noise about a more permanent solution,” I said. “At the very least, some bigwig oughta leap at the chance to save 3 cents of electricity a year by putting the complaint-spewing server out of its misery.”

I’d already resolved to contact Leila, the new head of HR, about this ticket. She needed to hear about Hothead, the manager-type who, in a fit of frustration, had taken a hair dryer to the same printer and nearly trashed it. I could also tell her about these lingering network vulnerabilities that would continue causing problems for everyone down the line.

I decided to do it without telling Reynaldo. I had the feeling that if I let him in on it, his eyes would roll clear out of his skull.


Working from Reynaldo’s cube, I assigned the HR printer a new IP address and once again cleared its queue. Then I sent Tony, the ticket holder, a private message asking him to give printing another go. I felt pretty good about his chances, but I’d wait for him to give me the all-clear before closing the ticket.

With the messaging app still open on my phone, I saw that Megan had invited me to drop by her cube. From Networking Central, I hoofed it up to Developers’ Row.

Megan’s cube was a familiar spot filled with bright, cartoony posters and figurines, the only splashes of color for miles. The titles and characters drew blanks in my over-the-hill brain. I kept meaning to ask her what they were, why she liked them. Another time, maybe. I found her with her back toward me, leaning intently in her swivel-chair toward the single monitor over her laptop docking station.

“Now a good time?” I asked under my breath.

She whirled around, tense, then relaxed with a smile, her gaze livelier than I’d seen it in a while. She picked up a small external memory stick from her desk to proffer my way. “That’s the HR application’s source code!” she murmured in conspiratorial fashion. “Guess where it came from?”

I smirked. “I know the answer should be ‘a code repository,’ but in this joint, that’s asking too much.”

“You’re right!” she replied. “It was never in a repo, never in source control. It lived and died on one dev’s local machine, a dev who retired way before I got here. His successor hung onto all the old code from that guy’s machine, just in case. But yeah, we don’t officially support the application anymore.”

Megan turned around, using her keyboard to tab over to an open window in a code editor. From over her shoulder, I saw line after line of an archaic programming language that the Egyptians might’ve used to build the pyramids.

“This code’s an undocumented disaster,” she said.

“There was an IP conflict at that,” I told her. “The print requests were going to a server that’s still running this thing. We assigned the printer a new IP, but the mystery server remains a head-scratcher.”

“I’ll trace through and figure out what it does when it gets a print request. Improve its error-handling in general,” Megan promised, more excited than daunted by the prospect. “I mean, if it is still up and running, I could tweak, recompile, and redeploy it so it doesn’t—”

Insistent metallic knocks sounded behind us. “Excuse me!”

It was the sort of pointed voice that somehow ignored your ears and stabbed into your gut instead. Megan froze, tense. I turned to find a woman at the threshold with the bearing of a vindictive hall monitor, rapping a fist against the cube’s bare metal frame.

She leveled a withering frown at Megan. “I was hoping for a status update on the Hewville refresh! Did I hear you say you were planning to work on something else?” The question sounded more like a threat.

This had to be Megan’s boss. Megan remained tense from head to foot. “I—”

“I need you to focus on your assigned projects. The things you can actually bill your time against.” After delivering the condescending reminder, the boss’ glare shifted my way. “And you are?”

I shoved unease aside and put on the game face I’d spent decades perfecting. “Tech Support. I meant no harm, ma’am. I was just asking Megan for help with an open support ticket.”

“She doesn’t have time for that!” the boss scolded. “If you truly need help from this department, then you must escalate your ticket through the proper channels.”

I already knew how that went: pulling together screenshots, logs, and other detailed information, only to receive half of a sentence fragment a few days later, asking for something I’d sent with the first message. In this case, I knew someone would just wag a finger at me about the HR application being out of support. Thanks but no thanks.

Megan struggled to muster one last defense. “This program’s still running on a server somewhere!”

“For changes to existing code, the proper procedure is to file a formal change request through the Project Management Team. The PMT will create a billing code and assign appropriate resources, if they deem it a proper use of development time.” The boss then looked at me like I was a used tissue she was ready to throw in the trash. “If that’s all, I suggest you head back to Tech Support, Mr. … ?”

No way was I handing her my name on a platter. I tried to glance Megan’s way, but she was staring at her lap. I felt bad leaving her in that lurch, but staying would only make things worse for the both of us.

“Later,” I said, both a goodbye and a promise.

I got the hell out of Devsville, slipping down flight after flight of stairs with a lead weight in my chest. For a moment, Megan had brightened in the face of a collaborative challenge. I’d felt a little more alive, too.

Thank God someone had been there to make sure no one helped each other.

The same sicknesses plagued the joint year after year. Almighty budgets. Status quo worship. Hierarchy and miles of red tape. Promising young people like Megan had their spirits crushed, and schmucks like me just put their heads down. Still a huge pile of other support tickets waiting for me, after all.

The frustration and resentment, the desire to do something to fix this, burned in my chest like a bonfire.

Back at my desk, I found a message from Tony confirming the new printer was behaving at last. Bolstered by my friends’ findings, I closed his ticket with notes about how the resolution was only a band-aid on a gaping wound. I messaged Megan, too, thanking her for trying.

And with that emotional bonfire still raging, I settled in and typed out a long email to Leila, detailing in full the most recent shenanigans I’d been a part of.

By the time I finished, I had one bit of good news: Aggie, my old mentor who’d turned manager a few years back, had accepted my meeting request to meet the next afternoon.

She wasn’t my boss, which meant it was still safe to vent her way. I had every intention. My resentment would no longer let itself be buried under this or that technical hiccup. It was insisting upon action.


The next day, I walked to the downtown coffee shop well ahead of the appointed time, glad to have the excuse to be somewhere else. Bought myself a drink and sat down where I could watch the door.

Five minutes late turned into ten … then twenty. No Aggie.

It was totally unlike her. Sure, she’d canceled last-minute before, but she’d always got in touch with me to let me know.

Caffeine jitters fueled a fear I couldn’t shake off. I sent PMs and left voicemails on her cell and work phones, asking her to respond when she could. Back at work, I asked around the department, including my boss.

She’d been AWOL the whole day. No one knew what was up.

Hours of work still stretched in front of me. I could barely sit in my chair, much less look at my monitor.

Just call me, Aggie, I willed, staring at the cell phone clutched in my hand.

She didn’t. Not that afternoon, not that evening. Beneath my fear was this strange gut feeling, this knowing sense that my worries were justified. It was crazy, but there was no talking myself out of it.

Early the next morning, they roped the whole department into the big conference room. Nobody knew what was going on until a small, ashen group of managers, directors, and directors of directors filed to the front and called for attention.

I already knew, deep down, what was coming.

“We’ve we received some very unfortunate news,” one of the senior directors spoke. “Agatha Shaw … passed away in her home after a heart attack.”

A few gasps escaped the assembly. Otherwise, you could’ve heard a pin drop. Wide-eyed looks of shock surged through us like lightning.

Aggie.

My gut had known all along, but my brain still wasn’t having it. Somebody somewhere must have goofed up royal, I thought. Happens all the time in this joint. Aggie had more life and fight in her than I ever did. I—

“Those of you who reported directly to Ms. Shaw will report to Bill Watson for now,” the director continued. “Dismissed.”

The brass began showing themselves out.

The rest of us just sat there, too stunned to move. That’s it? I marveled. No memorial? No counseling? No time off? Not one drip of sympathy?

I could only look on helplessly as Bill Watson, my boss, walked right over to me. He put his hand on my shoulder, leaned over, and muttered into my ear: “My office.”

While my brain reeled, my feet stood me up obediently. They marched me right off to the next chair I dropped into, the one opposite the large desk in Bill’s office.

He settled in on his side. “It’s a real shame about Aggie. We’re screwed without her.”

I just sat there, still reeling.

”My manager is looking for someone to step up in a big way.” Bill nodded in my direction. “You’re ready. You deserve it. Her direct reports are reporting to me for now; over the next few months, I’m transitioning them to you. The promotion will follow, as soon as the next performance review comes around!”

We were just cogs in that joint. Never before had the point been driven home so viscerally. Righteous rage surged up from within, clearing my brain and shooting strength through my limbs. I jumped out of that chair and glared down at him. “Hell no! Find someone else!”

I got the hell out of there and dragged myself all the way home to collapse on my couch.


To be continued ...

[Advertisement] ProGet’s got you covered with security and access controls on your NuGet feeds. Learn more.

Error'd: Einfach so

Do you say "a FAQ" or "an eff eh cue"? Peter says eff eh cue I think.

"This is a test" Peter G. harrumphed testily. "Create an FAQ with exactly nine entries. Nine? Nine."

67f8e4c8a1e640379d670a4039d1e54b

"I think I spent over $NaN?" said an anonymous. "This was an interesting offer on myminifactory.com with tight expiry date. Didn't claim."

9b22730c427144faa646895be1411d53

And a different reader expected a speedy delivery anon. "It was about 23:10 UTC when I took this screenshot, the time zone I keep my PC in, yet local time my pizza was estimated to arrive at 19:05 CST. Naturally, I should expect to receive my pizza about four hours ago! Not my typical experience with delivery as of late, I must say, but a welcome change nonetheless..."

cd43a695c8e846cf858cc3b72ba77094

Super saver Michael R. lamented "That hurts, I missed 6 coupons that would have saved me 0%."

ec9a5b21636b4b4bad633abdb4393820

And our dragoncoder047 ground this out between his teeth. "Refactored the runtime spritesheet packer in a game engine I contribute to, and wound up with this extremely helpful error message. Turns out that the problem was the ggggggggggggggg wasn't properly detecting ggggggggg and was putting all the ggggggggggggg's in the same ggggggggggggggggggg. I think. Ggggggggggggg!"

189885dc898f4fe4b7ef095e9f1ef99e

[Advertisement] Utilize BuildMaster to release your software with confidence, at the pace your business demands. Download today!

Flushed Out

While a project manager is frequently called upon for their planning ability, the real skill we want from project managers is their ability to communicate. The job of a project manager is to align the team doing the work, with the organization goals driving the work, with the management and leadership teams trying to understand the work, while juggling all the constraints like budgets, timelines, and the endlessly changing expectations for the project. A good project manager is worth their weight in gold. A bad one will cost their weight in gold.

Mark was hired on as a contractor, reporting to Tegan. Tegan was fresh out of business school, complete with an MBA and a variety of project-management training certifications. Unfortunately for Mark and the rest of the team, and especially unfortunately for Tegan, she had absolutely no real world experience. To make matters worse, this wasn't just a software project: they were working on a system which matched newly developed software with newly designed mechanics and custom build control electronics. A group of experienced software engineers, mechanical engineers, and electrical engineers all found themselves reporting to a bright and shiny MBA. It's a role that she probably could have grown into, but management saw all the acronyms she continuously put after her name, and decided she could just take the whole thing over with no real guidance.

It went badly pretty much from the beginning. Tegan was not a talented communicator. For example, Mark's team needed to know: on what timeline were the electrical engineers going to deliver the first prototypes, so the software team could start running bench tests of their software? Tegan's response was a fortune cookie message about balancing the complicated pipelines and lanes on the Gantt chart and hitting all of their milestones; like a fortune cookie, it was vague, important sounding, but ultimately empty.

Of course, the natural reaction amongst the engineers was to just route around the damage: the various teams could talk to each other just fine without going through Tegan. That, unfortunately, did not go over well with management. Tegan, as the project manager, was their insight into the project. They needed her in the loop on everything. And she couldn't just be informed, she had an MBA. She needed to be making decisions. But she was unqualified to make those decisions, which meant the project gradually ground to a halt. Tegan's emails got more vague, her meetings got longer but accomplished less, and after a certain point, she just stopped replying to key email threads.

The first few days of radio silence seemed like a gift. But as time passed and Tegan seemed uninterested or unable to reply to any of the questions the team had for her, the project started to flounder. The engineering teams escalated this problem to management. Management presumably went back to Tegan. At some point, feeling the weight of everything going wrong around her, Tegan sent out this email, which is definitely the best and clearest communication she managed during the project. It's arguably the clearest, and most accurate communication one could make in this situation:

Team,
I understand all the issues but there are complex interrelations that must be worked out. I am currently constipated on each issue and will let you know when there is movement.

  • Tegan
    MBA, CAPM, PMP

Her email cost the project many person-hours as all the engineering teams took a break to have a good laugh about the project manager admitting, in writing, that she was full of crap.

There was, eventually, movement. Tegan moved on to a new position at a different company. Her replacement, Pam, wasn't a new hire, but instead a transfer from another department. She wasn't a great project manager, certainly not worth her weight in gold, but she had enough experience to avoid the worst mistakes, and most important: she was good at regular communication in order to keep things moving.

[Advertisement] Keep the plebs out of prod. Restrict NuGet feed privileges with ProGet. Learn more.

CodeSOD: Module Test

TJ inherited a NestJS project. The original developers left the team many years ago, but they've left their mark in the codebase.

// ProjectsModule.ts
@Module({
  controllers: […],
  providers: […],
  exports: […],
})
export class ProjectsModule {}

NestJS is a dependency-injection oriented framework for TypeScript code. It offers "providers" (dependencies that can be injected), "controllers" (as one would expect), and lets you bundle them together into "modules". Modules can depend on other modules, letting you build a modular and flexible graph of dependencies. This means that the empty module isn't wrong here.

No, for it to be wrong, we need to write some tests:

// ProjectsModule.test.ts
describe("ProjectsModule", () => {
  it("can be created", () => {
    const projectsModule = new ProjectsModule()
    expect(projectsModule).toBeTruthy()
  })
})

Since modules are just containers for related code objects, there isn't much to test here. While "dynamic modules" which execute code are a thing, they don't execute that code at construction time anyway. This test will always pass. It isn't a test, it doesn't do anything. It likely doesn't even get their coverage up, since whatever providers or controllers it's referencing aren't covered by this test. It's a test that tests nothing but the framework it runs on top of.

"At least there are tests," TJ writes.

[Advertisement] Picking up NuGet is easy. Getting good at it takes time. Download our guide to learn the best practice of NuGet for the Enterprise.

CodeSOD: On Hold

"Dragoncoder" supports a web application that has a "wait time" for access. I hate that that's a thing, but I recognize that there are real-world constraints where this might make sense. Still, I hate it. But that's not the WTF.

      var minutes = parseInt( 12 , 10);
      var time = document.getElementById('waitTime');

      if ( minutes < 2) {
        time.innerText = "Your estimated wait time is 12 minute."
      } else if (minutes < 60) {
        time.innerText = "Your estimated wait time is 12 minutes."
      } else if (minutes === 60) {
        time.innerText = "Your estimated wait time is 0 hour."
      } else if (minutes < 120 && (minutes % 60 === 1)) {
        time.innerText = "Your estimated wait time is 0 hour and 12 minute."
      } else if (minutes < 120) {
        time.innerText = "Your estimated wait time is 0 hour and 12 minutes."
      } else if (minutes > (60 * 4)) {
        time.innerText = "Your estimated wait time is more than 4 hours."
      } else if (minutes % 60 === 0) {
        time.innerText = "Your estimated wait time is 0 hours."
      } else {
        time.innerText = "Your estimated wait time is 0 hours and 12 minutes."
      }

This wait time page is initially rendered by their backend, but after that point, gets served up by a cache at the edge of their CDN. That makes sense, since "have the users hammer your backend while they're waiting" is a bad idea.

Note the line var minutes = parseInt( 12 , 10);. This is rendered from the backend, which is of course my least favorite way to send data from the server side to the client side.

But that's not the core problem here. The core problem is: what the hell are they outputting?

If your wait time is less than 2, or less than 60, we tell you that your wait time is 12 minutes. Or "12 minute", because who cares about pluralization? If your wait time is exactly 60 minutes, we tell you that your wait time is "0 hour", which I assume means you'll have enough time to watch the classic airplane disaster movie, Zero Hour, which you surely know Airplane! is a remake of.

I can only think that the text is also being generated by logic on the server side- though our submitter doesn't suggest that's the case. Though they do wonder why the code couldn't be something like: Your estimated wait time is: ${Math.floor(wait_time_minutes / 60)} hours and ${wait_time_minutes % 60} minutes, which is both fewer bytes to send from your cache and more useful to the end user.

Or maybe we just make this wait time go away. Again, I don't know why it's there, there may be a good real-world constraint that requires it, but… is there? Is there really?

[Advertisement] Utilize BuildMaster to release your software with confidence, at the pace your business demands. Download today!

Best of…: Classic WTF: Difficult Personality

As the US took this weekend to celebrate their complicated relationship with tyranny, we reach back through the archives for another story of tyrants. If you think about it, the Declaration of Independence is basically the same thing as quitting without notice. Original --Remy

It was Steve's first week on the job, and he had plenty of questions about the code base and the new features he was supposed to implement. He muddled through for most of the week, but Friday morning he hit a brick wall and needed to talk to Bill, the architect.

"Can I meet with you for like an hour to go over things?" Steve asked.

"No."

"Can I get half an hour then? I h-"

"No. Company meeting, every Friday, 12-5pm. It should be on your calendar. I'll forward the invite."

Bill also couldn't free up time in the morning, so that meant Steve was stuck until Monday afternoon. Still, it probably wasn't all bad. He assumed that since this was a small company, in startup mode, it was going to be one of those meetings that was less meeting and more party. He had heard about one company in town that had a kegger every Friday afternoon.

Steve really should have known better. During his interview, the actual technical questions were thin on the ground. It focused more on "soft skills", like time management. He fielded a lot of questions about how best to manage his time. The other question that really stuck out in his mind was the standard, "Have you ever had to deal with a difficult personality in the workplace? How did you deal with it?" It was memorable, less because the question itself was unusual, but because at least six variations of the same question showed up in the interview.

On his very first day, he learned who the difficult personality was: Frank, the boss and grand-high pooba of the dev team. Around 2PM Frank lumberghed himself into Steve's cube. "Yeah, we've got a little problem," Frank said. "I've noticed you spending a great deal of time in the break room."

"Oh, yeah, I was just going back for more coffee," Steve said with an awkward laugh. "You know how it is with programmers- we're fueled by caffeine."

"Yeah, well, if you could just go ahead and make sure you're at your desk doing work, that would be great."

As it turned out, Frank had gone easy on Steve because it was Steve's first day. The next day, Steve sat in on Bill's planning meeting- a 4-hour marathon to organize the development backlog and parcel out work. Halfway through, Bill called for a break. He and a few other co-workers darted outside to gradually commit suicide via cancer, while everybody else hung around the room committing suicide by donut. And then Frank walked in.

"What's happening?"

"It's um… just a little break," one of the devs replied. "Bill's outside."

"I see." Frank loitered in the room until Bill returned. The instant Bill's foot crossed the threshold of the meeting room, Frank's human facade was stripped away, and a spitting, slavering demon replaced him. He proceeded to dress Bill down, back up and right back down for disrespecting his team, disrespecting the company, disrespecting Frank and Frank's poor elderly mother with his attitude. He closed with, "They're developers and I want them sitting around and developing! Not waiting for you to finish your smoke breaks!"

On Thursday, Steve got to drive a meeting to show off the latest batch of features the dev team had completed. When he turned on the projector, Frank asked, "What's wrong with your computer?"

"Um… nothing?"

"The desktop is wrong! None of the icons are in the right place!"

Like most developers, Steve had changed the wallpaper and reorganized his desktop to suit his working style. Unfortunately, his transgression against the default desktop settings set Frank off on a long rant that consumed the entirety of the meeting. Steve was lucky, Frank claimed, that he wasn't fired on the spot. Standard work was vitally important, and personalization was frowned upon. "It's vitally important that any developer can use any other developer's computer- we can't afford to waste a minute of time just because you needed to be a special little snowflake!"

By the time the Friday afternoon meeting rolled around, Steve should have been expecting some kind of Frank-led time management course. Instead, Bill handed him a mop. "New blood gets mop duties. Start in the break room, and then hit the other common areas."

"Excuse me?"

"Frank's orders. Every Friday, we spend the afternoon cleaning the office, from top to bottom."

"There isn't a cleaning crew?"

"Oh, there is," Bill said. "Frank doesn't trust them to do a good job." That weekend, Steve decided to take Frank's lessons on time management to heart, and immediately left to seek employment that didn't involve wasting his time.

[Advertisement] BuildMaster allows you to create a self-service release management platform that allows different teams to manage their applications. Explore how!

AI in Linux

The role of AI tools (LLMs, mainly) in Linux is under discussion, or it was, until Linus Torvalds “put his foot down” in support of the use of AI in Linux kernel development.

I can identify two major ways in which AI is used for Linux kernel development: authoring code and reviewing code. There are, at the time of writing, just over 1,200 kernel commits with an “Assisted-by” tag, from September 2025 to the present, most of which indicate patches which were written or assisted by LLM tools.

The second important use of AI for Linux comes with a new code review tool called Sashiko, which generates code reviews for patches considered for various subsystems. Sashiko ignited the current debate on AI in Linux because it pushes the envelope on AI in Linux: people who oppose or do not want to use AI could previously just refrain from using it to write their patches, but now there is a growing expectation that anyone who wants to contribute to Linux will have to interact with Sashiko or other AI tools like it to iterate on AI-generated feedback on their work.

One of the major lines of this discussion in the Linux kernel community has been with respect to the ethical considerations of the use of LLMs. Linus shuts this line of reasoning down entirely, firmly grounding the discussion in technical merits and rejecting any political discourse on the matter:

The kernel project has been and will continue to be about the technology.

Sure, the social angle of working on open source is important and often a very motivating part of the project, but in the end that’s a side benefit, not the point of the project.

This is NOT some kind of “social warrior” project, never has been, and never will be.

In the kernel community we do open source because it results in better technology, not because of religious reasons.

This argumentation is disingenuous and hypocritical. Linux is a political project and Linus is a political actor. Consider the use of the GPLv2 for licensing Linux. One can argue from technical merits – for instance, the copyleft nature of the GPLv2 pushes people, and in particular commercial entities, to upstream their drivers and other contributions into the Linux kernel. This contributes to the technical excellence for the kernel as a result.

But is this not a political choice, and a political act? The purpose of this choice is to influence the behavior of others and to advance the interests of the kernel ahead of their own. And Linus stuck to this decision for political reasons when GPLv3 was introduced, reasoning from morality and ethics when objecting to the license and the manner in which it was deployed, and called for a tacit boycott of the FSF.

Linus, and Linux, wields a tremendous degree of power and influence over the world, and it should be wielded responsibly. When Linus says the following:

Linux is not one of those anti-AI projects, and if somebody has issues with that, they can do the open-source thing and fork it.

Or just walk away.

I find it completely disingenuous. Linus is surely aware that, for all practical purposes, Linux cannot be forked. It is the world’s largest software project, and one of the most well-funded, too. The institutional knowledge among its contributors, the prospect of keeping up with the blistering pace of change, or even putting together a group of people with the time and funding to understand and maintain even a fraction of the kernel’s code independently of upstream, is, quite simply, intractable. Linus knows, this, too – it’s an explicitly cited reason for decisions like the use of the GPL, GPL-only symbols, and the unstable internal kernel ABI: to make the process of independently maintaining a fork of the kernel as difficult as possible.

People are right to petition Linux upstream to amend its behavior and policies before resorting to the impossible. Working on Linux requires practicing politics, both internally – see for example Linus’ response to the discussions around bcachefs, which had high technical excellence and low social/political competence – and at the intersection of Linux and the rest of the world. Therefore, we must table political, moral, and ethical arguments when we discuss how we go about the work. It’s a cheap, weak argument to direct the discussion away from political and ethical considerations when it wouldn’t serve your point and to table it when it would.

I’m willing to believe that the LLM-powered Sashiko code reviews provide a lot of good insights. However, to address just one externality of this tool, consider that AI companies are driving up the price of consumer hardware. An AI powered code review may improve a patch, but that patch won’t be of much use to the increasingly large cohort of people who are being priced out of the hardware they could run Linux on to enjoy the better patch.

Looking at it from another angle: how many tons of CO₂ added to the atmosphere or liters of fresh water supplies disrupted is a tolerable price for a better code review? Temperatures during heat waves are exceeding 50°C in India, causing tens thousands of deaths. The AI built-out is by far the fastest growing energy consumer in the world, and they’re being built with fossil fuels, or drawing green energy demand away from replacing the fossil fuels depended on by other industries. The same process is pricing regular people out of the energy they need to power their air conditioner during those heat waves and the technology it enables is pushing them out of the labor market and into poverty.

These externalities are very real, and are affecting a lot of people, including Linux kernel contributors and maintainers, and their friends and families, who are trying to bring these problems to Linus’ attention.

And what of the intangible effects of the use of these tools in Linux? Linux is lending some of the project’s immense influence towards legitimizing the makers of these tools, and Linus’ insistence on focusing narrowly on the technical applications of these tools is a generous gift to them. They will be sure to mention his support in boardrooms, meetings between lobbyists and governments, and anywhere else it will advance their interests.

It’s fair to ask: what are those interests, and what are they doing with this influence?

Photograph prominently featuring Mark Zuckerberg, Jeff Bezos and his wife, Sundar Pichai, and Elon Musk
Google CEO Sundar Pichai pictured at Donald Trump’s inauguration, together with other influential commercial leaders in AI. Sashiko primarily depends on Google Gemini for Linux kernel code reviews.

There are a lot of smart, passionate people who care about these kinds of questions. Where is the technical excellence in refusing to do this moral calculus, refusing to allow anyone else to do so, and driving away the talented Linux contributors who care about these problems? Linus Torvalds, the Linux community, and all of our communities should have the courage and insight to address these dimensions of the AI question honestly and in good faith.

Kylie Jenner Is Barely Using Her Own Meta Glasses

By: Nick Heer

Meta, last month:

We’re launching Meta Glasses with three frame styles offering distinct silhouettes that suit different faces, moods, and occasions:

[…]

  • Meta Glasses by Kylie — A unique slim oval frame designed in collaboration with Kylie Jenner and inspired by her personal style.

Katie Notopoulos, Business Insider, today:

Interestingly, since the launch of her namesake frames, Jenner doesn’t seem to be wearing them much in public. At the World Cup finals, she wore what appeared to be different tapered oval sunnies, and at a Knicks game, she used an old-school point-and-shoot camera to take photos of her boyfriend, rather than Meta glasses.

She did wear them in a recent Instagram post promoting her swimwear line.

Surely a good sign when a celebrity endorses a product they do not really care for. Imagine lending your entire image and your voice to something and then not promoting the hell out of it. Then again, imagine getting paid all that money and then realizing you are not contractually obligated to do so.

⌥ Permalink

Grain Is a Photo-First Social Network Built on AT Protocol

By: Nick Heer

Chad Miller has launched Grain, described as:

[…] a photo-first social app.

Post galleries, share 24-hour stories, and browse feeds by following, For You, camera, or location.

The most obvious comparison is Instagram, but without videos. I have been trying Grain for several months; it is really fun and worth checking out despite an icon that is, charitably, difficult to love. It is built on AT Protocol, the same as Bluesky, so you can log in with the same credentials. Here is a recent photo set I posted.

If you are on an iOS 27 beta build, you will probably need the TestFlight version that includes a fix for a crashing bug.

⌥ Permalink

Judge Dismisses Google’s Lawsuit Against SerpApi

By: Nick Heer

Barry Schwartz, Search Engine Roundtable:

Last December, Google sued SerpApi over scraping its search results, and now a court has granted SerpApi’s motion to dismiss the case. U.S. District Judge Yvonne Gonzalez Rogers dismissed these claims with leave to amend, giving Google 21 days to refile its complaint if it can demonstrate authorization from copyright owners.

You can see the court filing here (PDF) which basically says Google brought claims under Section 1201 of the Digital Millennium Copyright Act (DMCA), alleging SerpApi bypassed its anti-bot barrier (“SearchGuard”). However, Section 1201 only protects technological measures that restrict access to copyrighted works.

I had complicated feelings about this lawsuit because while I think there is little distinction between the web scraping activities of each party, and the effect of limiting SerpApi would be to reinforce Google’s illegal search monopoly, a win for SerpApi also seems to invalidate the meagre control website owners have over scraping. These restrictions are noticeable during normal web use and are quite frustrating. But without effective copyright reform, there really are few options for choosing whether you want your work to be incorporated into training data.

Reddit’s comparable lawsuit against SerpApi, among others, is ongoing.

⌥ Permalink

Google Says It Still Sends ‘Billions of Clicks’ to Websites

By: Nick Heer

Google’s Nick Fox on, for some reason, LinkedIn:

People also asked if AI in Search would mean that people never click through to websites anymore.

Actually, as we’ve shared before, we continue to send billions of clicks to the web every day through Search. And we’ve designed our AI features in Search to connect people to websites. In fact, we’re now sending billions of clicks to websites every week through AI features in Search alone – and we’re just getting started.

The specific effects of Google’s A.I. search features is made murkier by all the other features the company has to answer queries immediately instead of sending them to an external website. Nevertheless, the claim that Google still sends “billions of clicks” is just another way of saying Google still has a monopoly on the world’s web searches. It is a substantially similar statement to the one Google issued a year ago, and something the company is only repeating because the New York Times asked for comment.

If only Google had some kind of dashboard or console with information about a website’s search performance.

Barry Schwartz, Search Engine Roundtable:

Google even built out the AI performance reports in Search Console but excluded clicks. This is because Google is “continuing to work with website owners to understand what insights will be most helpful to inform their strategies” and Google doesn’t think those are clicks. Nope, because Google does not want us to see our click-through rates from AI Search features compared to traditional search features.

I am sure this is a mere oversight that Google will be eager to correct.

⌥ Permalink

Gurman: Apple Set to Launch Device Leasing Program

By: Nick Heer

Joe Rossignol, MacRumors:

Apple and Klarna are partnering on a new “Apple Upgrade” program set to launch in the U.S. on Tuesday, July 28, according to Bloomberg’s Mark Gurman.

The program will allow you to finance most iPhone, iPad, Mac, and Apple Watch models, with a 24-month term for iPhones and Apple Watches and a 36-month term for iPads and Macs. Customers will be able to pay off the device early during the term, upgrade early to a newer device, or keep or return the existing device after the term.

In 2017, working with McKinsey, Apple said its supply chain would eventually become a closed loop, without specifying a timeframe or even a firm methodology for how it might do so. Perhaps encouraging people to treat devices as leased objects exchanged every few years gets closer to this goal, as Apple can capture a greater number of sold devices. (Be honest: how many of you have a bunch of old products sitting around unused? There is gold in them thar hills.)

Then again, the other thing I thought about is the rising cost of components, and the negative effect increasing prices could have on new purchases. Spreading the cost out over monthly payments might make some people feel less burdened by a dramatically pricier new Mac — something like what has happened with car sales. In the United States, for example, the most common way to purchase a new car is increasingly through financing, which divides the purchase price into monthly payments with added interest. Because the cost is divided up, it means manufacturers can increase prices and buyers can be pulled into longer and costlier payment plans.

Apple, of course, already has financing options for its new product purchases — Affirm in Canada, and a slew of options in the U.S. — as it has rapidly become a bank. If a large enough number of new Apple product buyers choose financing instead of outright purchases, it could incentivize more expensive products and longer payment plans. If you read that and got excited, you probably own lots of Apple stock.

⌥ Permalink

Blaming Data Centres, Alberta Electricity Industry Expects Tripling of Power Prices

By: Nick Heer

If you believe Meta, its forthcoming Albertan data centre will be, at worst, unnoticeable to people across this province:

We pay the full costs of our data centers’ energy use so consumers aren’t negatively impacted, and fund new and upgraded infrastructure. We worked closely with Greenlight Limited Partnership, Altalink, Capitol Power, and the Alberta Electric System Operator to plan for and meet our energy needs years in advance of this data center coming online.

Sure sounds like the commitments of a stand-up corporate citizen. However, if you listen to the power companies, they tell a somewhat different story.

Rory White, the National Observer:

In its latest earnings call, [Capital Power’s] CEO Avik Dey predicted a “return to higher pricing” telling analysts that he couldn’t rule out prices of “$80 or $90” per megawatt hour of electricity by early 2028. By 2029, rival company TransAlta is predicting an average of $100 — more than triple the average price this year — a high not seen since the province’s 2021-2023 energy crisis.

Driving this increase is the province’s reliance on natural gas power, which cannot be built fast enough to accommodate data centre demand, according to Will Noel, a senior electricity analyst at the Pembina Institute. “In Canada and internationally, we’re seeing this huge crunch in gas turbine supply shortages,” he said. The province has also backed itself into a corner in terms of its options for alternatives. “Alberta has over the past couple of years really stifled the growth in its wind and solar.”

White reports this would increase the electricity costs to a typical household by “hundreds of dollars more per year”. However, as White caveats, these prices have precedent as recently as three years ago — and that undersells it. In 2022, electricity costs were as much as five times above current rates. Even without adjusting for inflation, we are currently paying less than we did from 2018–2020, and a tripling of the current rate would have a nominal cost similar to that of 2008. That is not to say this is good, but I think it is worthwhile seeing the “tripling” figure within the context of recent fluctuations.

⌥ Permalink

Apple to Recap WWDC Announcements in Calgary This Week

By: Nick Heer

An un-bylined report from Calgary.tech:

Apple will make a rare conference appearance in Calgary next week, joining the inaugural Swift Rockies gathering for iOS developers at the Calgary Zoo.

Taking place July 22 and 23, Swift Rockies is a boutique, single-track conference created by Calgary-based iOS engineer Raman Singh. The independently organized event is capped at 180 attendees and designed to encourage closer interaction between speakers and developers through round-table seating and an intimate format.

Via Jason Anthony Guy, who writes:

[…] Even if the session being presented is one they offer globally, it’s fantastic to see Apple’s Developer Relations team again relating to developers outside of an Apple-managed event. That’s a welcome shift from the last few years that I was there (even pre-pandemic). My favorite part of being on that team was direct developer engagement. I hope it marks a return to form for WWDR.

Swift Rockies is sold out. According to the conference’s website, the presentation is open to any registered Apple developer regardless of whether they have a pass, though it does seem to be largely a recap of the biggest WWDC announcements. I have no idea if there is still space but, if you would like to go, you need to submit your request to Apple by midnight.

Also, if you are going to Swift Rockies and have never been to Calgary, send me a message and I can give you a couple of recommendations for things to do near the conference, if you want.

⌥ Permalink

Meta Faces Tennessee Trial Over Instagram Design

By: Nick Heer

Diana Novak Jones, Reuters:

Meta Platforms faces trial in Tennessee on Monday over the state’s claims that Instagram’s design is to blame for a youth mental-health ​crisis, one of several trials in the coming weeks testing allegations that the company’s social media platforms were intentionally built to be addictive.

I imagine Meta will spend considerable time trying to nail down a definition of “addiction”. The lawsuit makes no attempt to define the term, though it does quote several internal Meta communications acknowledging this as an outcome of the company’s products.1 In a legal sense, this might be an important question.

The term, however, seems loaded: is Instagram really comparable to cigarettes? As a matter of practicality and ethics, I have personally found it more useful to think of this in terms of whether Instagram and its competitors are intended to be hypnotic and compelling beyond users’ comfort levels. I believe they are. A social media app today looks more like gambling, where you wager your time, than it does a continuation of real-world social experiences.

It is therefore too bad this lawsuit and others like it are exclusively related to the effects they have on children. I understand why, but I think we all need greater control over what we see. There are people I know whose TikTok feeds are absolutely full of A.I.-generated nonsense, often a mix of not-harmful trash and malicious information. This will not be solved by simply telling people of all ages to just say no to using these apps.


  1. If the name “Skrmetti” rings a bell for you, it might be because he was a party in one the U.S Supreme Court decisions that permits discrimination against trans people↥︎

⌥ Permalink

Copyright Is Not Enough

By: Nick Heer

Zoey Forbes, a U.K.-based attorney, in the Dial:

One hundred and forty years on from the Berne Convention, the basic principles of copyright remain the same, but GenAI poses a threat to authors different from anything that has existed before. Its novel technology is not only destabilizing what it means to reproduce works, but what it means to produce them.

This is a terrific and well-rounded exploration of copyright law and generative artificial intelligence from a non-U.S. perspective. That matters because the U.S. has a carve-out for “fair use” of copyrighted works, which is something generative A.I. companies are relying on for their defence of their unethical and maybe illegal practices. If it holds, it makes the rest of the world less desirable for A.I. development which, I fear, means it becomes a race to the bottom for all the countries that want a slice of this well-funded pie.

⌥ Permalink

Breach of A.I. Music Generator Suno Reveals Scraping of Music From Deezer and YouTube

By: Nick Heer

Jason Koebler, 404 Media:

The AI music generation tool Suno scraped millions of songs and lyrics from YouTube Music, Deezer, and Genius, as well as from the stock music libraries Pond5, Jamendo, Freesound, the International Music Score Library Project, and podcasts via RSS feeds, according to a hacker who breached the company and shared data about Suno’s training libraries with 404 Media. The hacker was also able to access user information for hundreds of thousands of Suno’s customers, as well as Stripe payment information, they said.

Suno is fighting several lawsuits, including one filed by UMG in which it makes the argument its use is sufficiently transformative. But people have been prosecuted for the mere act of downloading hundreds to thousands of songs. This line of argument suggests to me that, beyond some astronomical number of downloads, it becomes entirely legal because no single song will be of much consequence.

⌥ Permalink

Permanent Daylight Saving Time

By: Nick Heer

Dr. Drang:

A couple of days ago, Casey Liss took a break from arguing about temperature scales to tweak me about the recent passage of the Sunshine Protection Act by the House. The Act would make Daylight Saving Time permanent, something Casey knows I disapprove of. A similar bill passed the Senate a few years ago, and Donald Trump has said he will sign this one, so there’s a decent chance it’ll become law. Let’s see what will happen if it does.

Like many people, I was all ready to abolish DST until 2013, when I read Drang’s article advocating for the twice-yearly ritual of clock changing. I was converted.

Five years ago, the Albertan government asked voters whether we should “adopt year-round Daylight Saving Time, which is summer hours”. A bare majority, 50.2%, voted against it. So, of course, our government has adopted permanent DST, and we will not be turning our clocks back this November. This aligns with the practices of our neighbouring provinces.

The United States at least has the advantage of a large population living fairly far south. On the shortest day of the year, Los Angeles still sees nearly 10 hours of daylight and over 14 on the longest day. In somewhere as far north as Calgary, the difference in daylight hours is far greater — from under 8 hours in December to over 16 in June. That means the effects of permanent DST are highly acute. The sun will not rise here before 8:00 am from October 15 through February 8, with the latest sunrises at 9:39 am for several days in a row.

Theory is different from reality and perhaps my mind will be changed if it is still daylight after 5:00 pm on the shortest days of the year. I look forward to that. But this has been tried before — unsuccessfully — and I question why this time would be any different.

⌥ Permalink

And Another Thing About Squircles

By: Nick Heer

If you cannot get enough squircle talk, Michael Tsai has a whole roundup of different takes, some of which you may not have seen already.

Here is another thing: robbing icons of a variable for creative excellence and flexibility is just another form of context collapse. You know how Instagram and YouTube are host to professionals and amateurs alike? There are benefits to creating an impression of similar legitimacy, but the rigidity of social media platforms’ formats also collapses the difference between legitimate information and absolute nonsense.

Losing some of the artistry in an icon makes it more difficult to distinguish between legitimate and well-crafted Mac apps, and everything else. It is not a perfect proxy, to be sure, and there are some very nice squircle icons. But it is nevertheless an unfortunate change that makes apps of varying levels of quality look more similar.

⌥ Permalink

Amazon Will Print Book-Length A.I. Gibberish on Demand

By: Nick Heer

Kashmir Hill, of the New York Times (gift link), learned about an A.I.-generated biography of her for sale on Amazon for $27. It is just one of many wholly-generated books available there:

Amazon does not mind if people hawk A.I.-generated books on its platform, unless they are truly and deeply terrible. “Charlie Kirk: An Inspiring Journey of Young Political Conservative and Activist Who Fights for America,” published in February 2025, became an Amazon best seller after Mr. Kirk was killed last September — which means it probably sold thousands of copies. But after dozens of scathing reviews called it “mind-numbing,” “a scam” and “a disgrace,” Amazon took it down.

Too bad about all the trees killed for this print-on-demand nonsense, though at least it fits with Amazon’s decline into one of the world’s biggest retailers of sketchy, counterfeit, and knock-off products.

⌥ Permalink

Morgan Stanley Says the ‘Space’ in SpaceX Is Worth About $8 Per Share

By: Nick Heer

Bryce Elder, on the Financial Times’ Alphaville blog:

Here’s paragraph one of the Morgan Stanley’s SpaceX initiation note:

With an ‘X of 1’ position in space infrastructure, we believe SpaceX can convert energy into intelligence at scale with optionality to monetize through a range of consumer and enterprise solutions for the next era of AI… the final frontier.

[Adam] Jonas — formerly the Wall Street bank’s Tesla and occasionally other auto companies analyst — was last year reallocated to the sort of free-radical futurologist role we’d thought had been cornered by Shingy.

Today’s news that SpaceX has already started dipping below its IPO valuation reminded me of this piece. Remind me — is being compared to Shingy good?

At any rate, Morgan Stanley estimates that the combination of X and Grok will grow this year to generate only a little less revenue than Twitter did in 2021, its last full fiscal year as a public company. Fear not, investors, as the bank also says SpaceX will make $17 billion in “enterprise A.I.”, and then nearly triple that next year. I guess everyone wants to give money to the CSAM and misogynist fantasy generator for business.

⌥ Permalink

Apple Updates Its Advertising Policies

By: Nick Heer

Sarah Perez, TechCrunch:

In a newly published Apple Advertising Services policy, effective as of July 14, 2026, the iPhone maker shares its rules for advertising on Apple Maps. Notably, it prohibits the broad category of home services businesses, like plumbing, electrical, locksmith, HVAC, pest control, roofing, and general contracting services, among others.

If Apple is interested in updating its advertising policy further, I suggest none. I spent lots of money on its nominally premium products, and I pay a monthly fee to use the company’s services. Alas, here we are.

Apple is also prohibiting ads for cryptocurrency ATMs. Also, there is this:

2.7.2 Ad content that directly or indirectly promotes or facilitates the sale of products or services that compete with Apple hardware products (e.g., mobile phones/smartphones, tablets, notebook/desktop computers, and smart watches) is reviewed on a case-by-case basis.

I do not see this in its previous restricted or unacceptable ad guidelines.

Eric Benjamin Seufert, who writes Mobile Dev Memo, noticed the phrase “on the relevant Apple software applications or Apple devices” has been changed to remove references to Apple-specific devices or software:

The new language could simply accommodate the availability of Apple-owned services on the web and through third-party devices and operating systems; the Apple TV app, for instance, is available on smart TVs, streaming devices, and game consoles. But the addition of “other properties” is conspicuously broad and appears to give Apple the contractual latitude to distribute ads beyond its own services entirely. This would allow for a material expansion of the company’s advertising surface area.

Apple Maps is also available on the web and is used by DuckDuckGo, so this could simply be covering that inevitable broader placement. But the reality is that Apple is now operating an advertising business, and putting more ads in more places is one way to make the numbers go up. Not the user satisfaction numbers, of course — I cannot imagine a single user who wants more ads in their life, unless they own shares in this company. But it will print money at no cost, and that is what the people running a huge corporation with little competition want to do. It was not inevitable until Apple made it so by opening the door.

⌥ Permalink

The Shape of Apps

By: Nick Heer

Paul Kafasis, writing on the Rogue Amoeba blog:

With last year’s release of MacOS 26 (Tahoe), Apple made a mess of app icons. In the first betas of MacOS 27 (Golden Gate), however, there are signs of a turnaround. We’re urging Apple to continue making improvements, by restoring the ability for MacOS app icons to have distinct shapes.

Kafasis reignited my simmering frustration with the mandated squircle in MacOS. My Dock contains three of the apps shown in the collection in this post: MarsEdit, NetNewsWire, and Sketch. I like their current more-uniform icons fine enough, but they are less distinguished than the ones these applications used to have.

But maybe that is the whole point?

Louie Mantia, writing on the Parakeet blog:

The shape of apps is a squircle. And it has been proven to work for everyone. Companies can use their logo as an app icon. Designers can create something specifically for the platform. And both of these get to look like an app. Whether people consider the squircle a container or a canvas, this uniform appearance communicates its function: a squircle represents an app, just like how a piece of paper represents a digital document, or a folder represents, well, a folder.

Semantically, there’s something really beautiful about that. As an icon designer, I appreciate that different types of things have visual distinction.

Mantia touches on all the pragmatic reasons to unify the shape of icons in an operating system, all of which I have considered, and then drops the above paragraphs — and things started to make more sense to me. This is a different way to think about it. This is not a situation where a cleaner looks like a zesty beverage with toxic consequences. The shared general function of these icons does help communicate something and makes them less ambiguous in that sense.

But a broad category of functionality is only part of the story of an icon and — with respect to Mantia’s long and illustrious history of work in this area, and that of co-Parakeeter Luka Grafera — taking away a difference of shape also limits what an icon can communicate. It may not be a cleaning product that looks and is packaged like juice, but imagine if every consumable liquid was in identical bottles with only a different label. You might go for something refreshing after a workout and end up drinking soup. Sure, you can argue the label for a beverage should not look similar to the one for soup, but the two would be far easier to distinguish if they were not in the same package.

I still think constraining designers to a singular shape, while more clearly defining specific objects as apps, has made it harder to distinguish between them, particularly when combined with the glassy and contrast-killing layer effects of Tahoe. Happily, though not retreating on the squircle, the shapes within icons in MacOS Golden Gate are at least more clearly defined.

⌥ Permalink

Meta’s ‘Activity from Other Businesses’

By: Nick Heer

Yash Garg:

A friend pointed me to this option in Meta Accounts Center called “Activity from other businesses”, which shows “activity sent from other businesses or organizations to show you relevant content.” Their main help page doesn’t even work in my region.

If you dig around in your Meta account privacy settings, you can turn this feature off. While you are there, though, you might take a scroll through the audience-based advertising list. These are companies that have, most often for me, “uploaded or used a list” of email addresses or phone numbers. In my case, this is hundreds of businesses — some of which I recognize or can see why they would have my contact information, and many of which I have never heard of. Advertisers are supposed to have permission to use this information, of course, but there is basically no way for me to confirm whether I gave permission or report that I did not. Restricting further use is also comically unfriendly: it is a button two levels deep and, for all but the first four advertisers, you must click a “See more” button each time to display the full list.

Update: Rodrigo Ghedin counts four separate privacy-hostile things Meta has done or is rumoured to be working on since the beginning of June.

⌥ Permalink

Lies Told About Data Centres

By: Nick Heer

Karl Bode:

One lie that companies have been telling local municipalities is that if they greenlight a massive local AI data center, it will immediately bring a flood of savvy innovators to your podunk-ass town.

The promotional materials for Kevin O’Leary’s still hypothetical “data centre park” imagine a futuristic campus full of bright young minds doing complicated A.I. stuff on-site in north-central Alberta. But why would they be there — fifty kilometres from the nearest city and 500 kilometres from Edmonton, the nearest major city? Why would they not be in Vancouver, or Silicon Valley, or anywhere else with an internet connection?

Also:

These companies aren’t coincidentally aiming construction at states and municipalities that are too broken and corrupted to put up meaningful regulatory opposition. […]

That is one reason why we are seeing a bunch of these proposals here. And these companies benefit from a lack of transparency.

⌥ Permalink

⌥ Stories Told About Data Centres

By: Nick Heer

Nathaniel Rich, author of the novel “Cloudthief”, in a non-fiction retelling for New York Times Magazine of a 2007 heist of a London data centre by Terry Ellis and others:

The fixer — Ellis called him Ray and won’t reveal his name — met him in North London near Hampstead Heath for coffee and cakes. When it came time to discuss business, to avoid being overheard, they strolled into the park.

Ray had brought Ellis a few jobs before. But this job, he warned, was of an entirely different order. As Ellis claims in “The Art of Robbery,” a self-published memoir written after his release from prison, he eventually learned that Ray had been contacted by a consultant employed by “some influential bankers from America.” The bankers “were involved in prime mortgages” and had “circumnavigated” certain regulations. Damning evidence of these circumnavigations could be found in banking files held in the King’s Cross area in a giant building known as a data center.

This is a dramatic story, and one I think should be read with a heavy dose of skepticism. It seems that most of the criminal details have been shared by Ellis. For a start, the claim that some bankers ostensibly contracted with “Ray” is just a little too perfect for a recession-era tale. These bankers are pretty much universally loathed, and this justification makes this theft seem more palatable than a simple financial motive. For example, there was a similar data centre theft in October 2006, which would be unrelated to the lending crisis in the following years.

Another problem is that Rich says crimes like these are covered-up in part by a data centre operator because they are loathe to “admit to flaws in its security, [which] would only encourage additional attacks and scare away its clients”. Therefore, the lack of evidence for the specific circumstances of this crime is supposed to be a buttress for its likelihood, not a weakness, which is not reassuring.

The story of the theft was, as far as I can tell, broken by Here is the City, then a gossipy financial news site:

The data center itself is thought to be used by a number of companies, including JPMorgan, which is believed to have told staff that some of its systems could be off-line for parts of the day today as a result of the theft. Fortunately the thieves are thought to have got away with just the computer hardware, and not any sensitive information which may also have been stored at the facility.

Tom Espiner, of ZDNet, a few days later:

Reports circulating on the Internet last week that JPMorgan, a customer of Verizon Business, had been affected by the burglary were incorrect, according to a source at the investment bank. There has been no loss of service or data, said the source.

On the one hand, of course all these parties tried to cover this up. The reading-between-the-lines story implied by these early reports and Rich’s telling is that some banking higher-ups, perhaps from JPMorgan, wanted to cover up some crimes, and denying any meaningful effect is just more cover-up. But little of this is substantiated by contemporary or current reporting — which is, of course, the whole problem with using a lack of evidence as the foundation for a story.

Rich, in the Times:

“The banks knew they were sending mortgages to people who couldn’t pay back,” he says today. “That’s what broke the whole system. That was the big con.” Ellis remains convinced that the bankers who paid for the Verizon job wanted to destroy evidence of their involvement in fraudulent subprime mortgages — the inside information that Ellis received about the data center, he believes, “would have had to come from the top” — but he can’t prove it. He never saw what was on the servers.

“Our job was to get the motherboards,” he says. “We were paid quite handsomely. Whatever happened after that was none of our concern.”

In contemporaneous reports, the Metropolitan Police noted the theft of motherboards and processors. But if these bankers wanted to cover up their fraudulent practices, surely the hard drives would have been the target, right? In Rich’s version, entire servers were taken, so perhaps this is just a misunderstanding.

This story smells fishy. I believe the theft happened, of course, and Ellis’ involvement, but I am not as convinced this had anything to do with covering up some white collar crime. (By the way, the Guardian in 2018 published an interview with Ellis about the interesting prison where he was transferred and which led to his rehabilitation.)

The heist element is only about half of Rich’s story; much of it is a discussion about data centre secrecy:

The public fogginess about data centers is not an accident. It is the product of a willful strategy by the world’s largest tech corporations, whose business models rest on the public assumption that the internet, and all the data it holds, is as immaterial as air — or as a cloud, to borrow the metaphor commonly used to describe the sum of information stored on servers. As the digital-media scholar Tung-Hui Hu writes in “A Prehistory of the Cloud,” the cloud “hides its physical location by design.”

[…]

It was a lot easier to defend data when people didn’t know it existed. The more people learn about data centers, the more they hate them. […]

If you read a website like this one, you were probably aware that data centres were commonplace twenty or more years ago. Like the one near King’s Cross, some were hidden in plain sight, while others were purpose-built facilities that look like hangars stuffed with servers. But the A.I. boom has meant rapid increases in the speed, scale, and quantity of data centres. People quickly learned not only of their existence, but how much pressure they put on local resources. Tech companies, it seemed, were caught by surprise; and as someone who spends a lot of time immersed in this world, so was I.

Much of the consternation I have seen in more general audiences has been about data centres in general. People simply were not aware that Amazon has warehouses full of products, and other warehouses full of computers. As Rich writes, this is deliberate, for business secrecy reasons, security, and environmental costs. But, also, I think some of that unawareness is because of just how boring it is. If nobody wants to know hidden information, is it really a secret? It only became one when the information these companies were hiding had real-life effects.

It does seem that public awareness is putting pressure on corporations to improve data centres and make them more efficient. But that is not a standard. New data centres are powered by petroleum with a pinky promise of renewable offsets. In some regressive regions, like Alberta, new power plants for data centres must be powered by methane gas. In a further complication, Meta’s proposed data centre is scheduled to be completed before the power plant is ready, meaning it will be dependent on existing grid power for perhaps years. Meta’s is just one of the data centres proposed for Alberta. Another one, a gigawatt cluster, would also require a dedicated gas-fired power plant, while Kevin O’Leary’s questionable project is supposed to require over three times the combined power of those other two.

For years, the tech industry told us we did not need to have much concern for how digital products and services worked, and many of us did not bother to find out. But it turns out the demands of our email and Netflix subscription were comparatively easy to hide. At the very least, what we ought to demand from projects with the scale and ambition of these data centres is open disclosure of their power consumption, water use, and emissions.

But we ought to demand more than the bare minimum. Transparency does as much good as a big banner reading we are destroying the planet but we are also creating a lot of value for shareholders. When a single data centre is projected to use about as much power as the entire city of Calgary is currently — Enmax says 1,260 megawatts as of writing — we should have a say in whether that makes sense. A.I. remains a thing that is happening to us rather than with or for us. It is built on assuming consent and asking forgiveness, which has more-or-less worked for the industry and gave it way too much confidence. Tech companies could have spent decades being better corporate citizens. Data centres are just one part, but they are representative of the difference between the stories told by tech companies and the things we can actually know.

Perhaps Meta Should Not Have Spent Decades Being Creepy

By: Nick Heer

Meta, in a press release called “Meta’s A.I. Glasses: Your Questions Answered”:

Can’t people just cover up or disable the LED?

The camera is disabled when people try to do this. Beginning with our second generation of glasses, the camera is automatically disabled if we detect that the capture LED has been blocked. No photos or videos can be taken until we detect that the light is unblocked.

Since the introduction of this safeguard, we’ve seen some people go beyond using tape to sophisticated efforts to modify or destroy the capture LED. We are continuously improving our ability to detect tampering, and now we’re updating the glasses to disable the camera if they detect the LED was physically tampered with or destroyed. No other kind of camera has done this and we’re proud to lead the industry forward.

Meta is not being entirely honest here. For many, many years, Apple’s laptops have contained a camera indicator light with among the highest security protections possible. Over ten years ago, iSight cameras were not adequately secured. I cannot find a more recent example showing a similar vulnerability, at least suggesting a better level of protection in today’s cameras. Other computers also have built-in cameras with in-use indicators, with various approaches to security, some of which have vulnerabilities. In general, though, it is not new for an indicator light to resist tampering as long as the camera remains functional.

What is different is the threat. Cameras built into computers need protection mostly from remote attacks, while Meta’s glasses need protection from an owner deliberately altering them. Meta is selling creep glasses and hoping it can outsmart everyone buying them — and that the rest of us similarly trust Meta to protect our privacy.

David Gerard, Pivot to A.I.:

But it gets better! Meta’s planning a new version of the glasses that records continuously! [FT, archive]

a new hardware line of smart glasses that would continuously record audio while taking photos every few seconds.

And they won’t have a recording light. […]

This product is still rumoured, and perhaps the shipping version will have some kind of external recording indicator. But given the way Meta would intend a device like this to be used, I doubt it will, otherwise it would be indicating basically all the time.

Meta is in the business of asking for forgiveness instead of seeking permission. It will release these regardless of public approval, and with the belief it can control how that continuous recording is used. But determined people will surely find workarounds and vulnerabilities, and Meta surely believes it can trust itself to fix them. But I do not. Meta has not earned the right to ship anything like this without incurring deep suspicion about the product and anyone using it.

⌥ Permalink

Maybe assigning TCP connection to Linux traffic control 'flows'

By: cks

Back when I wrote about using tc to limit the outgoing bandwidth of a web server, I expressed a wish:

For our purposes it would be nice to do something sophisticated to aggregate all HTTPS requests from a single IP address together in a single 'flow' for tc-sfq(8). In theory this is possible with tc-flow(8) but in practice I can't figure out a command line that works right (despite consulting eg tc-sfb(8)'s example).

The good news is that I managed to figure out a syntax that tc would accept. The bad news is that I don't know if it does what I want, because I'm not sure how you see what connections are assigned to what flows.

I'll put forward two variations of tc-flow(8) commands. To start with, I'll set the stage. We start with our bandwidth limited class that will have a tc-sfq(8) qdisc below it, straight from the first entry:

tc class add dev eno1 parent 1: classid 1:10 htb rate 400mbit ceil 500mbit prio 10

Our first option is to do exactly what I said above, making the flow per-IP instead of per four-tuple, and attach this flow classification to our bandwidth limited qdisc.

tc filter add dev eno1 protocol ip parent 1:10 handle 20 flow hash keys src,dst,proto,proto-src divisor 4096

(The 'handle <something>' is apparently critical although I don't know why. I got this wisdom from here after extensive online searches.)

This theoretically makes flows be based on the source IP, destination IP, source port, and the protocol. In our case, for a HTTPS web server with us dealing with outgoing traffic, the source, protocol, and source port are all the same, so this only varies the flow by the destination IP. This is what we want to aggregate all connections from the same IP. Some sources will tell you to use 'perturb N' on this. We don't want to do this because we want flows to stay sticky and not get re-sorted every so often, and we want to set a very large divisor so that we maintain a lot of distinct flows to go with our (initially) large number of connections.

(The tc-sfb(8) manual page has an example like this, basing flows purely on the destination IP, but it's attached to the root qdisc.)

The other option is to be even more aggressive and theoretically group flows by the /24 of the destination IP address, and nothing else (since we're already theoretically only dealing with HTTPS responses sent by our web server, with the protocol, source IP, and source port all constant). This is done with the other sort of flow filter, a 'map' filter instead of a 'hash':

tc filter add dev eno1 protocol ip parent 1:10 handle 20 flow map key dst rshift 8 divisor 4096

If I'm understanding tc-flow(8) correctly, this takes the destination IP and right shifts it 8 bits (then divides the result by 4096, once again to try to have as many distinct flows as we could in theory have simultaneous connections). Dropping the right octet of a full IP address effectively reduces it to its /24. I think we could also use 'dst and 0xffffff00' to get basically the same effect, but maybe that would get more flow collisions when the dust settled.

I believe that we also need to tell our tc-sfq(8) queue how many flows it's supposed to have, matched with the number we used above:

tc qdisc add dev eno1 parent 1:10 handle 10: sfq flows 4096 divisor 4096

Unfortunately, I don't know if all of this works because I don't know of any way to see what flow a given connection is classified into (either before or after tc-sfq(8) gets at it). It could be that my tc-flow(8) usage isn't actually assigning flows the way I want (especially in the second case). It could also be that I need to put my flow filter somewhere else in the tc class hierarchy, or maybe somehow attach it directly to the tc-sfq(8) qdisc.

(Given the example in tc-sfb(8), where the filter is attached directly to the sfb qdisc, possibly I need to attach the flow filters to the sfq qdisc, with 'parent 10:' instead of 'parent 1:10'. This appears to be accepted by tc, although who knows if it's working.)

PS: Based on what 'iftop' is telling me, I'm relatively sure that my flow filter isn't actually working the way I want (with a flow filter either in this version or directly attached to the sfq qdisc). But it's hard to be sure.

The story of how we (eventually) found our missing firewall rule

By: cks

I recently shared a war story about how we had load problems on a new web server but not an older one because of a missing firewall rule. In a comment, Aristotle Pagaltzis asked how we'd realized that we had a missing firewall rule. This is a good question, because the sequence of events involves a certain amount of luck and coincidence. So here's that story, which may be a useful example of system administration in action in practice (ie, it's kind of messy).

When we first deployed our new web server for a highly in-demand data set, Apache wasn't configured for very many concurrent connections and was immediately overloaded with HTTP requests, but we mostly shrugged. Because this upset our monitoring system and I did want to monitor the web server at least a bit, I turned up the Apache connection limits to an absurd number. This was mostly enough, and the web server's outgoing bandwidth jumped up to 1G wire rates, which at the time I thought was fine. Soon after that, an apparently unrelated machine began to have NFS performance problems, and after some work we realized that it was on the same 1G switch as the highly active web server and its network traffic was getting crowded out.

In a rush to limit the web server's bandwidth usage so the much more important other machine would stop having NFS problems, I hastily added some tc based bandwidth limits by hand, and by this I mean that I typed 'tc' commands in a shell session. A few days later, we worked out how to make mod_qos based limits work in Apache itself. One of the limits we imposed and fiddled with was a limit on the number of concurrent connections from a single IP address, which had a much bigger effect than I expected. The Apache error log showed that some IPs were hitting it quite a lot, and our metrics system said the number of concurrent Apache requests dropped dramatically afterward (to about half what they'd been before).

After the dust settled from the immediate crisis, we needed to decide if we were going to make the tc-based limits a permanent part of the machine's configuration (making it the first machine officially using tc in production, and we'd have to write something to install them on system startup) or if we'd rely entirely on mod_qos. While considering the tradeoffs involved, I remembered that we had a general purpose 'block brute force things' system on our perimeter firewall, so we should probably make it apply to HTTP and HTTPS requests to our new web server for extra insurance. Immediately after I did this, our metrics system showed another major drop in concurrent Apache connections (shrinking by half again).

(The perimeter firewall's system has a list of IP addresses that it applies the HTTP and HTTPS limits to, and the new web server, with a new IP address, wasn't in the list (until I added it).)

That major drop from the firewall rule was what sparked my realization of what was different that would explain why our main web server mostly hadn't been overwhelmed by this traffic, because that was the only thing out of all of our rate limiting changes that our main web server had in place. Our main web server had no tc limits and we'd turned off mod_qos several years ago and never revisited that change.

PS: We decided to keep both the Apache mod_qos limits and the tc based limits, partly for extra insurance. We've decided that we really don't want this particular web server to run over the bandwidth limits, so having two mechanisms that limits it means they'd both have to fail.

(We're not yet worried enough to look into FreeBSD PF's features for bandwidth limits, partly because if we got it wrong on the firewall, we could affect a lot more than this machine.)

Making sense of diskless workstations through two models of them

By: cks

When I wrote about how early SunOS did diskless workstations, I got a good question in a comment:

Obviously this is slow, but having an indeterminate number of users sharing a circa 1982 mechanical hard drive sounds like a problematic amount of slow. Was this ever truly worth doing? I know hard drives were wildly expensive back then, but surely the work slowdown from having multiple users sharing a single hard drive in this fashion would mean the ROI for those hard drives would seem obvious?

One answer is that diskless workstations were not infrequently used in situations where there was no 'ROI' as such, for example for use by university graduate students (generally not in dedicated offices, unless you were very lucky, but instead in shared in terminal rooms). But in my view, a deeper answer is that there are two usage models of diskless workstations.

In one model, a diskless workstation was an inferior substitute for a workstation with a local disk. It had to do its disk IO over a slow shared 10MBit network connection to a server with mechanical HDDs that were used by multiple people (on all of those diskless workstations), which was obviously much worse performance than a local disk. But you saved the cost of the local disk, and perhaps you bought a low end workstation model (such as the basic Sun 3/50 instead of the better, faster 3/60) because they weren't going to be fast anyway.

In the other model, a diskless workstation was a superior replacement for a serial terminal, in much the same way that X terminals were later. Instead of a single text 'window' and no local computing, you gave people something with multiple windows, graphical capabilities, and some degree of local computing that could be faster than an (over)loaded central server. In the process you might save money on the central server, since it didn't need as much compute capacity as it would if everyone was directly logging in to it (although diskless workstations cost a lot more than serial terminals, so you probably weren't saving money overall).

(The other advantage of the diskless workstation model over the serial terminal model was it was more amenable to incremental upgrades, since you could buy better workstations (perhaps with disks) one by one. It was a "personal computer" model instead of a "terminal" model.)

These two models lead to different calculations of costs and benefits. In the first model, you're losing productivity but saving money on hardware, and the question is how much does the lost productivity actually cost you. You're probably going to give diskless workstations to people who don't have highly valuable productivity. Diskless workstations are a downgrade and workstations with disks will become a status symbol, a sign that you're important enough to call for the extra expense.

In the second model you're gaining productivity at the cost of spending more on hardware. the question is how much extra productivity do people gain compared to the extra cost of diskless workstations over serial terminals (possibly factoring in a cheaper server, fewer serial lines and serial port boards in the server, and so on). In some cases the productivity gains may be significant at relatively modest extra cost. Diskless workstations are an upgrade and a status symbol (compared to serial terminals).

Of course these models cross over somewhere, so you get to look at the relative payoffs and costs of serial terminals, diskless workstations, and workstations with local disks for different groups of people with different productivity payoffs. Once NFS and other shared, writable filesystems entered the picture, things got more complicated because even your 'local disk' workstations might be NFS mounting home directories, shared work areas, and so on, both for collaboration and so that people weren't tied to specific physical workstations.

(When X terminals arrived they added another section to the spectrum, with graphics but without local computation. This could make sense in a variety of ways; some people benefited from graphics but not local computation, and some people needed to do most or all of their compute on the same machine as their data (on HDDs), instead of hauling it back and forth over shared 10MBit networks with the network filesystem protocol of your choice.)

There's likely also a practical commercial aspect to diskless workstations. My impression is that Sun and other early Unix workstation vendors were relatively desperate to get their machines into places (Unix was a new thing, after all), so in the grand tradition of such things they created a low cost entry level version of their product to get their foot in the door, even if it wasn't all that great. If you could initially sell a company some base configuration diskless workstations and a server to go with them for cheap, maybe you could turn that into a later sale of better, more expensive hardware once the company got a taste of Unix.

(Sun would later continue this tradition by selling entry level hardware without hardware floating point.)

Ubuntu 26.04 has broken shutdown announcements and wall doesn't work

By: cks

Today, for reasons outside the scope of this entry, we needed to do unscheduled reboots on a number of Ubuntu 26.04 servers that people log in to and use. As is our usual process, we didn't reboot these on the spot; instead we ran 'shutdown -r +NN "<a message about the situation>"' so that people would have a little bit of warning because the impending shutdown would be periodically announced (by systemd, because this is systemd-based these days). Then, to our unpleasant surprise, we discovered that no announcements were happening. Shortly afterward we discovered that the venerable 'wall' program wasn't making announcements either.

Surprisingly, these turn out to be two separate issues. The wall issue is because starting in Debian 13 ('Trixie') and Ubuntu 25.10, the Debian and Ubuntu systemd is built without support for /var/run/utmp (aka /run/utmp), the traditional file recording who is logged in where; wall (which comes from the 'bsdutils' package) only looks in the utmp file. If there's no utmp file, wall is never going to do anything. If you need a wall equivalent, you'll need to write a script that gets the list of active user sessions with ptys and writes a message to them itself.

(For Debian Trixie dropping support for utmp, see eg this debian-devel thread. Apparently one reason for the change is that the utmp format has Y2038 problems. A replacement is available through the wtmpdb package and project, which also gives you a working 'last' command. You have to hook it up in your PAM configuration, and the Ubuntu 26.04 OpenSSH is built without wtmpdb support, so I believe you're going to be missing some information.)

The issue with shutdown not broadcasting messages appears to be because in Ubuntu 26.04 (with systemd 259.5), systemd's logind doesn't know what ttys people's SSH logins are using. You can see this with 'loginctl' or 'loginctl -j', which will have no TTY information for all SSH logins (although if you log in on the console, it will have that). This isn't the case in Ubuntu 24.04 (with systemd 255.4) or Fedora 43 (with systemd 258.9), although both versions are built with UTMP support (in theory this shouldn't matter, since logind's tty tracking is a separate thing). Logind's announcements of impending shutdowns only go to TTYs that it knows about, so since it doesn't know about any SSH login ptys, none of them get any announcements.

Update: Now that I pay attention, Fedora 43's systemd is older than Ubuntu 26.04's. However, Fedora 44 has systemd 259.7 (with UTMP enabled) and its 'loginctl' doesn't have this problem.

If you're using 'who' or 'w' on Ubuntu 26.04 (or at least the version of 'who' from GNU Coreutils, since the uutils version currently suffers from bug #2152801), you might notice that they do report pty information, at least if you have AppArmor disabled (as we do):

; who
cks      sshd pts/0   Jul 20 21:47 (...)
; lsb_release -r
Release:        26.04

This is because while the GNU Coreutils version of 'who' talks to systemd to try to get this information (through a set of systemd library APIs), if there's no TTY information for a session it will also look through /dev/pts to try to find a likely candidate. This works often enough that 'who' typically shows information for most interactive SSH sessions with a pty. The 'w' program, which comes from procps, has a similar fallback if systemd and utmp are both not reporting the tty. Since 'who' and 'w' use different approaches, on Ubuntu 26.04 one may report a tty that the other doesn't.

(Apparently the default Ubuntu 26.04 AppArmor profiles block access by 'who' to /run/systemd/sessions, which the systemd library API uses under the covers to get session information. Amusingly, this only affects 'who', not 'w', as 'w' has no specific AppArmor profile.)

Changes in how something behaves are a signal (but can be hard to notice)

By: cks

I mentioned recently that we'd moved a very popular data set from our main web server to a new one that only handled that data set. On the new web server, we found it necessary to set an absurdly high Apache connection limit of 4,000 concurrent requests, because Apache could run out otherwise (and could even run out at 4,000, it was just less often).

Our main web server has a much lower concurrent connection limit but it didn't experience these problems (and not because we'd imposed connection limits in Apache; we'd turned that off in December of 2022). But as we (I) wrangled with the new web server to get it to stop running out of connections and so on, and even after I got mod_qos working on it, I never paused to ask myself why the new web server had so many problems with this when the main web server had run for years without explosions (or at least very infrequent ones).

(Part of this was because the main web server had started exploding sometimes; that was why we'd moved this data set to its own server. But even those explosions had been less severe than what I was seeing.)

The answer is that for years, our perimeter firewall has had per-IP connection rate limits for HTTP and HTTPS connection to our main web server (among other per-IP connection rate limits, for example for SSH connections). These predate the modern popularity of this data set and were added to stop other abuse, but it turns out that a lot of the connection volume for this data set was coming from a few IPs that were opening up a ton of rapid-fire and often simultaneous connections. Once we applied per-IP limits on both the number of simultaneous connections you could have (in mod_qos) and the rate at which you could make connections (in the firewall), the new web server's connection count dropped like a stone (but people still kept pulling data from it as fast as they could).

In retrospect, the change in the web server's behavior when we moved this data set to a new host was a signal. We'd had only occasional problems on the old web server host one (despite it being actively used for other things) and we had constant ones on the new server (dedicated only to the data set). So we could have asked what was different, and then investigated, and then found the perimeter firewall issue. But on the other hand, this is sort of hindsight bias speaking. Such changes in behavior are a signal, but as system administrators we're drowning in signals and we have to sort out what's meaningful and what's either a coincidence or a consequence of something else (for example, a sudden increase in demand for this data set, which would have also explained why we were suddenly seeing problems even on the main web server).

PS: There's some recent evidence that there was a real but temporary shift in demand for this data set, in large part from people (or software) that make extremely inefficient requests. If these people are done now, or have improved their software, that would be nice. Perhaps all of the connection blocking and limited bandwidth have encouraged them to download things only once and then keep local caches.

The Rust coreutils (uutils) are sticky in Ubuntu 26.04 LTS

By: cks

GNU Coreutils are what they sound like; a GNU version of a bunch of basic, core Unix programs such as 'mkdir', 'head', 'chmod', 'cp', 'mv', and so on. For a long time, talking about 'GNU Coreutils' was unnecessary and you could just talk about 'Coreutils'. Then some people decided to rewrite Coreutils in Rust, the uutils coreutils, which still wouldn't be very important for most people except that Canonical decided to make the Rust versions the default in Ubuntu 25.10 and then 26.04.

In theory the Rust coreutils aim for 100% compatibility with GNU Coreutils and anything to the contrary is a bug. In practice, I found an incompatibility almost immediately when testing 26.04 pre-release, Ubuntu bug reports are useless, and I was pretty certain that other people on our systems would run into other issues, so I decided that we would sit out this round of Canonical making 26.04 LTS people mandatory beta-testers of their current passion project.

Canonical doesn't make it easy to switch from Rust uutils to GNU Coreutils, but they do at least make it possible. The magic apt-get command you need is:

apt-get install coreutils-from-gnu coreutils-from-uutils- --allow-remove-essential

(Taken from here.)

If you do this, I suggest that you immediately do 'apt-mark hold coreutils-from-uutils' (you may not want to hold 'coreutils-from-gnu', since there might be bugfix updates to it for some reason, although the real programs are in the 'gnu-coreutils' package).

The reason you might want to do this is, well, let me quote a Fediverse post of mine:

Ubuntu: you can totally continue to use GNU Coreutils in 26.04 LTS.
Also Ubuntu: build-essential depends on the new Rust coreutils.

Yeah, that's not "you can totally continue to use GNU Coreutils", although it sure is tempting to build my own build-essential package with a different dependency.

What this means in practice is that if you install coreutils-from-gnu, the build-essential package is uninstalled if you have it installed, and if you install build-essential later, coreutils-from-gnu is uninstalled and coreutils-from-uutils (the Rust version) is reinstalled. If you're not paying close attention to all of the messages that an 'apt-get' is printing out, you might miss this (especially if the apt-get is happening in the middle of your general install framework). Then you will be surprised, as I was, when it turns out that your 26.04 systems have Rust coreutils despite you theoretically having switched.

Build-essential itself doesn't do much, although installing it is a convenient way to get some core software building tools (especially if you want to build Ubuntu packages, perhaps to make local changes). What really matters is that 'apt-get build-dep' will insist on installing build-essential, and you may want to do 'apt-get build-dep <some package>' for all sorts of reasons. For example, if you're going to build your own Emacs, 'apt-get build-dep emacs' is a convenient way to get most or all of the development packages it's going to want, rather than looking them up and getting each one yourself.

The build-essential dependency is explicit:

$ apt-cache show build-essential
[...]
Depends: libc6-dev | libc-dev, gcc (>= 4:14.2), g++ (>= 4:14.2), make, dpkg-dev (>= 1.22.11), coreutils-from-uutils

Not 'coreutils' (a meta-package that depends on either), not explicitly 'coreutils-from-uutils | coreutils-from-gnu', a direct, specific dependency on the Rust coreutils. This turns out to be a Canonical bodge from September 2025 that's not in the upstream package (via). Since this is an explicit dependency, dealing with it requires things like building your own version of build-essential that has a fixed dependency (perhaps with dgit).

There may be other packages with specific dependencies on 'coreutils-from-uutils', which is why I suggested you explicitly 'apt-mark hold' it. With the package held (or both coreutils-from-* packages held), 'apt-get install <some package>' will abort rather than flip your Coreutils setup around. Then at least you can find out which new package will make you unhappy with Canonical.

("apt-cache rdepends coreutils-from-uutils' doesn't show me anything on our Ubuntu 26.04 LTS machines, but I don't know if that's complete across the entire Ubuntu package set for 26.04.)

How my desktops wound up with multiple D-Bus user session instances

By: cks

I mentioned recently (when I dug into systemd and your user D-Bus session bus) that some of my machines were set up so that they sometimes started another D-Bus session bus daemon for me. After having done some experimentation I can say that this isn't necessary (or really, proper) and I've now stopped doing it. You might wonder how I got myself into this situation, and that's a story of history.

On my primary desktops, I've never used any graphical login manager like gdm, xdm, or so on (which has long been sort of a heresy), and I've never run a standard desktop; instead I have my own window manager environment. This leaves me having to do a lot of things myself as part of starting the X server and my environment that a standard desktop and login environment takes care of for you.

D-Bus started being a thing in Linux before systemd. In those days, starting your user D-Bus session daemon was part of the jobs of your desktop environment, either internally or through files it put into the standard /etc/X11/xinit/xinitrc.d (where your graphical login manager of choice should pick them up, although I'm not sure how that works these days). Since I didn't have a desktop environment and was doing it all myself, I had to research what was normally run on session startup and duplicate it in my own shell scripts, and one of those things was running dbus-launch with the appropriate arguments. In the way I ran it, dbus-launch unconditionally starts a D-Bus session daemon and sets '$DBUS_SESSION_BUS_ADDRESS' to point to it.

This was fine in the pre-systemd days, when my regular console login didn't have a D-Bus session daemon started for it (or set up to be ready to start). Well, it was mostly fine, because the D-Bus session bus address was in /tmp, and things can happen to files in /tmp under various circumstances. But I think it only very rarely went wrong, enough that I didn't really notice.

When systemd started providing a D-Bus setup of its own, the proper official /etc/X11/xinit/xinitrc.d was changed so that it detected this and stepped out of the way and not started another D-Bus daemon (and desktop environments that did it all themselves internally were changed similarly). But my own scripts never had this in and never noticed, so when I logged in on the console systemd would set up the whole D-Bus stuff for me even though I was logging in on the console and then my xinit based scripts would promptly start another D-Bus daemon and override that.

(Well, systemd set up all of this provided that my session was in the right class.)

All of this shows one of the challenges of having your own desktop environment; it's on you to keep up with this sort of stuff, and you're probably not hooked into the information channels for it (as far as I know, the various desktop environment people talk to each other). There are probably other places where my environment has drifted away from how it should be.

PS: It's possible that I'll run into problems with my switch, because there's one potentially important thing that's different between the two approaches. The old approach started the D-Bus session daemon with a fully initialized environment (since I was starting it from my login shell after logging in), while the new one starts it with whatever minimal environment it gets from 'systemd --user'.

Argc and argv in early Research Unix

By: cks

Recently I was peripherally involved in a Fediverse discussion about (C's) argc and argv (the arguments to your main(), the start of a C program). Famously, argv[] is an array of pointers to your program's arguments (including the nominal name of the program), and it's sort of traditional to terminate it with a NULL pointer (although this isn't required by the Single Unix Standard; its execve() specification is silent on this). If you think about it, having both argc and a NULL-terminated argv is redundant, since you could determine one from the other. So me being me, I wondered how far back argc and argv went in Unix (and if argv was NULL terminated from the beginning). The answer turns out to be that they go all the way back to Research Unix V1, which is before C existed, and argv[] wasn't originally NULL terminated.

Update: Tony Finch pointed out that POSIX actually does specifically require argv[] to be NULL terminated (and the NULL not be counted in argc). See the comment for details.

The V1 exec(2) manual page is specific about both sides of the V1 exec() API (which is expressed in assembly language terms, since C wasn't invented yet). Exec() is called with a NULL-terminated array of pointers to the (zero-terminated) argument strings, but the invoked program receives an explicit count of the arguments along with an array of argument pointers, and the array is not listed as NULL-terminated. The V1 kernel source code for sysexec (in u2.s) doesn't appear to put in a final NULL pointer or any other pointer value after the regular argv[] pointers, so your program has to use argc to know when to stop.

The logic of this split between the exec() API and the API to programs is a bit clearer in the C code of the V4 exec() in sys/ken/sys1.c. Exec() needs to count the number of arguments in order to do things like allocate the correct size of argv[] array on the stack of the new program, and having created that count it might as well pass that to the new program as argc. However, if I'm reading the V4 exec() correctly, it adds a final '-1' right after the normal end of the argv[] array:

while(na--) {
  suword(ap=+2, c);
  do
    subyte(c++, *cp);
  while(*cp++);
}
suword(ap+2, -1);

This trailing -1 remains present all the way through the V6 exec() in sys/ken/sys1.c (and I don't know why it was -1 instead of 0; the V6 crt0.s doesn't seem to make any visible check for it).

Finally, in the V7 exece() in sys/sys.1, we get an actual NULL pointer at the end of argv[]. However, this is less of a terminator and more of a separator, because the addition of environment variables in V7 has turned argv[] into two arrays of pointers stacked on top of each other, one for the arguments and one for the user environment (which is also terminated with a NULL, because otherwise there's no way to tell). Based on how the C program startup libc/csu/crt0.s has a loop, I think that it finds the environment by walking the argv[] array to find the separator NULL, although the kernel is still providing argc as well as the argv[] array.

As far as I can tell, both System III and 4.2 BSD continue to add the separator NULL (it's more obvious in the 4.x BSD source, where there's an explicit copy of '0' into the user stack; in System III, it appears that the user stack section is pre-zeroed so the code just bumps the offset). BSD continued doing this at least as late as 4.3 BSD Reno (cf). Based on this repository, it appears that System V Release 2 for the Vax also separated argv[] and the environment with a NULL (cf vax/os/exec.c).

If there were Unix systems that later changed this to not have a separating NULL between argv[] and the environment (and thus not giving argv[] a terminating NULL), I don't know what they are. Instead, I suspect that either some C compilers on early non-Unix systems omitted the NULL at the end of their (made up) argv or that the ANSI C and POSIX people didn't want to explicitly require it.

Update: See above, POSIX does explicitly require NULL termination.

(Now you know why I was looking at exec() in early Unix and came to understand its argv size limit.)

The early Research Unix exec(2) argv size limit

By: cks

When I wrote up how V7 gave us environment variables, I mentioned that up to V6, exec(2) had a limit of 510 bytes of command line arguments (including argv[0], the nominal name of your program). You can see the check in the V6 kernel exec() code in sys/ken/sys1.c (where it returns E2BIG in this case). You might wonder where this limit comes from and why.

When you exec() something, you discard your current process's memory and address space to create create a new one for the new program. Your current (user) memory includes the argv you're passing to exec(), so the kernel has to copy it from your user space into the kernel and then back, temporarily holding it in some sort of kernel memory. In a modern kernel you might dynamically allocate this kernel memory in exec() through the kernel equivalent of malloc(), but the Research Unix kernels were simple and didn't have that sort of thing. Instead, through Research Unix V6, they got their temporary scratch space for exec() by allocating a disk buffer, reusing a facility the kernel already needed. These disk buffers were 512 bytes long, which is more or less where the 510 byte limit on argument size comes from.

(I don't know why it's 510 bytes instead of 512; I've been unable to follow the code closely enough to see if it slips in a use of the last two bytes of the buffer for something else.)

You might innocently think that using a disk buffer just pushes the problem of dynamic allocation of (kernel) memory back one layer, to the disk buffer system. However, early Research Unix kernels are more brute force than that. The V6 kernel has a fixed (and limited) chunk of memory reserved for disk buffers, the buffers array in sys/dmr/bio.c, with its size set by NBUF in sys/param.h. The default NBUF isn't very large, but early Research Unix ran on small systems and had low limits in general (the same param.h sets a limit of 50 processes for the entire system).

This straightforward approach to exec() and disk buffers goes back to at least Research Unix V4 (I haven't looked earlier than that). In V7, the kernel implementation is rather more complex because it needs to handle environment variables too, but it still sort of uses the disk buffer trick. In order to get the extra space without using much extra RAM, V7 uses swap space, writing to it and reading back from it through V7's general disk buffer system (which probably often meant that the disk buffer you wrote to swap is still in RAM when you read it back shortly afterward as part of setting up the new process's memory). So to copy the exec() and exece() arguments, V7 allocates a disk buffer in swap space, copies from user space to the disk buffer until it fills up, flushes and releases the disk buffer, gets a new disk buffer for the new block of swap, and does it all over again.

(As an extra complication, V7 didn't have page based swapping, it only swapped whole programs. So during an execve(), V7 allocated swap space for a NCARGS sized 'program' and then used as much of it as necessary, one disk buffer at a time. If V7 couldn't allocate the necessary swap space during exece(), it paniced.)

PS: If you look at the V6 code for exec() carefully, you'll see one spot where it does 'suword(ap=+2, c);', which looks odd and wrong. That's because in V6 C, '=+' was how you wrote in-place arithmetic, instead of '+=' in V7 and the C we know today.

Systemd and your user D-Bus session bus

By: cks

These days, a lot of things want you to have a (user) D-Bus session bus (to go with the system-wide one), which will listen for connections on some socket (the bus address). On a systemd based system, the normal D-Bus user session bus socket is /run/user/<uid>/bus, which you can see by logging in and doing, for example, 'echo $DBUS_SESSION_BUS_ADDRESS'. Suppose that you SSH in to some Linux machine that uses systemd, run this, see your expected D-Bus session bus address, and even use 'lsof' to see what's listening to it. Do you actually have a D-Bus session bus active?

Well, maybe, because these days your D-Bus session bus is a socket-activated systemd service. Specifically, it's a user socket and service that's managed by your per-user systemd instance, which is a systemd process running as 'systemd --user' under your uid. When this 'systemd --user' process starts, typically on your first login (including a SSH login), it will start listening on a bunch of sockets, including your standard D-Bus session bus socket, and it will insert a $DBUS_SESSION_BUS_ADDRESS into the 'systemd (user) service manager environment variables' with 'systemctl --user set-environment ...', where many things will then pull it back out (typically including your new SSH login).

Your actual D-Bus session bus and associated processes are only started by systemd if something actually tries to talk to the session bus. This will typically start 'dbus.service', but what that does varies from distribution to distribution. On Fedora, this runs dbus-broker-launch via /etc/systemd/user/dbus.service (which is actually a symlink to /usr/lib/systemd/user/dbus-broker.service), which I believe then starts dbus-daemon itself; on Ubuntu and Debian this directly runs dbus-daemon via /usr/lib/systemd/user/dbus.service. Your distance will likely vary on other distributions.

(The corollary to this is that 'systemctl --user set-environment' and friends aren't using your D-Bus session bus, because systemd does this without starting the the session bus. In Ubuntu 26.04 this communication is done through a systemd socket, /run/user/<uid>/systemd/private, that apparently uses a private API.)

If you SSH in to a server, look at your processes, and there's no dbus-daemon, I believe that you can be pretty sure that you don't have a D-Bus session bus operating yet. One corollary of this is that any surprising delays in logging in and starting your session definitely aren't D-Bus session bus problems, because clearly your session bus hasn't even been started.

(Well, assuming that you can deliberately start your session bus, for example by running 'dbus-monitor --session'. If your session bus refuses to start, that might be your problem. In my case, it's not.)

(This is the kind of thing that I want to write down in case I ever need it again, because I got confused about what was providing my session bus and whether it was activated or not.)

PS: Normally your session bus is shared between all logins, both on the local console and remotely over SSH, but this isn't required. It's possible to start another D-Bus session bus daemon and arrange for that session bus to be used by other processes. However I'm not sure you actually want to do this and it may be a mistake for some of my machines to (still) be set up this way.

Prometheus 3.14's (likely) duration functions, especially step()

By: cks

An exciting change landed in the development version of Prometheus recently, making PromQL arithmetic expressions in time durations a standard feature instead of an experimental one. For me, duration arithmetic expressions by themselves aren't the truly interesting part. What's really exciting is that as part of this change, Prometheus has added some new PromQL functions, especially step().

(Although step() is gated behind an 'experimental PromQL functions' feature flag today, it will be made available as a standard function as part of duration expressions becoming a standard feature.)

For people who are familiar with Grafana, step() is the PromQL version of Grafana's $__interval interpolation variable. When you're in a range query, step() is how big the range step is, which was previously unavailable in PromQL even though Prometheus obviously knew this information. Prometheus has also added range(), which gives you the full size of the range duration (and the *_of() functions for further selection if, for example, the step() might be too small). Since PromQL now allows arithmetic expressions in durations, you can use step() in them, allowing you to write PromQL expressions like 'rate(your_metric[step()])', where the duration will automatically adjust to whatever the range step is.

Where step() is especially handy for me is when I'm doing ad-hoc graphs directly in Prometheus's web query interface (instead of wrestling with Grafana Explore). Previously I had to go through various increasingly elaborate processes to find out what the step value was for a given time range, so I could plug it into rate() or various *_over_time() things or the like. Now I can just ask for 'rate(...[step()])' and it will all work out right, and it will keep working right as I zoom the time scale in or out.

In theory one could replace various Grafana uses of $__interval in PromQL queries with step(). In practice this probably isn't worth it, unless you're running into problems with $__interval for some reason or maybe when you're writing completely new queries for new dashboards and so on. The minor advantage of using step() even in Grafana is that you can easily copy the query out of Grafana and put it directly into Prometheus or a query tool to see exactly what you're getting.

To use step() and friends in Prometheus 3.13, you need to enable some feature flags. But since this is going to be in the next version of Prometheus (unless something goes wrong and the Prometheus developers have to back out this change), I think it's pretty safe to turn on the necessary feature flags and start using step() and friends now, at least in ad-hoc poking around. Even if you have to stop using step() later, it will improve your experience today.

(This elaborates on a Fediverse post of mine.)

How early SunOS did diskless workstations before NFS

By: cks

Over on the Fediverse, I had a little exchange recently:

[other person in a conversation]: I still haven’t forgiven Sun for NFS. No I’m not bitter.

@cks: It could have been worse, Sun could have stuck with nd.

What I was referencing in my post is a now obscure piece of cursed knowledge that I'm happy to share with you today.

Sun's workstations could boot without a local disk from very early on (because that made them cheaper, not because it made them better), but famously NFS only appeared in SunOS 2.0 (which required Sun to also create the idea of a virtual filesystem switch (VFS), which has appeared in basically every Unix since). The pre-NFS versions of SunOS operated without a local disk by using Sun's 'nd', the 'net(work) disk', which is basically what it sounds like.

SunOS nd(4) was a kernel block device (well, pseudo-device) that did its block IO through the network to the server kernel. As covered in nd(4), the same driver was used on both the client and the server, and the server handles everything in the kernel; nd(8) is only there for server setup purposes. The client wasn't configured with the server's information; instead it found the server through the simple approach of "[it] finds the server by broadcasting the initial request". As you can see from the fact that the manual pages I've linked to are for SunOS 3.0, SunOS kept the nd driver and infrastructure quite a long time after NFS was available (I'm not sure, but it might have only been dropped in SunOS 4).

(You can also see the SunOS 1.0 nd(4).)

By itself, the idea of a network disk device isn't particularly cursed. We used to run iSCSI based fileservers quite happily, there's a general ATA over Ethernet protocol that was at least a lot simpler than iSCSI, and Linux has DRBD (and there's probably others out there). What makes SunOS nd into something special is how it works, which comes from a specific limitation of SunOS covered in nd(4):

One last type of unit is provided for use by the server. These are called local units and are named /dev/ndl∗. The Sun physical disk sector 0 label only provides a limited number of partitions per physical disk (eight). Since this number is small and these partitions have somewhat fixed meanings, the nd driver itself has a subpartitioning capability built-in. This allows the large server physical disk partition (e.g. /dev/xy0g ) to be broken up into any number of diskless client partitions.

What this meant in practice was that on your server, you set up one giant partition and then manually decided on the starting and ending sectors for every nd 'disk' within that partition. Keeping track of all of these and making sure that they didn't overlap was your problem; as the nd(8) manual page dryly notes in the BUGS section, 'no sanity checking of disk partitions is done'.

For extra bonus problems, you might run out of available partitions to use on your server disk because you needed all of the available ones for regular filesystems and your swap area. If you were in this situation you could take the dangerous but necessary step of specifying your network disks using the special 'c' partition (cf dkinfo(8)), which was conventionally used to provide access to the entire disk. This was extra dangerous because you had to make sure that the nd disks you specified weren't overlapping into any regular partitions that you were using, since as nd(8) says, nd itself did no sanity checking. If you said sectors X to Y were network disk X, that's what they were, and goodness help you if some of them were also something else.

(I think this meant you could expose the server's /usr disk partition as a read-only 'public' nd device, so all your diskless clients could mount it, rather than having to put a separate copy of /usr into your nd area. This seems to be explicit considered in nd(4).)

Another charming thing about nd was that it didn't use UDP (or TCP). Instead it uses its own IP datagram protocol because, as covered in nd(4), "IP datagrams were chosen instead of UDP datagrams because only the IP header is checksummed, not the entire packet as in UDP" (and the manual page straight up says it was also done because the kernel internal interfaces were simpler). What this means is that the data sent through nd had no checksum protections on the wire, not even UDP's basic one; you were very much counting on absolutely nothing going wrong on your early 1980s Ethernet network.

(With early 1980s CPUs and so on, it presumably made a real performance difference to not checksum 1024 bytes of data on each packet. Early Sun workstations were not exactly performance powerhouses.)

All of this made using nd extra exciting and somewhat cursed. I don't think anyone really liked it back in the days, and people were happy to move to NFS, which used regular server filesystems and had much less of a chance to blow up your server and your clients in exciting ways.

Finding an outdated Git mirror host

By: cks

Suppose, not hypothetically, that you have a situation where there's a number of distinct hosts backing some Git repository, such as all of the IP addresses of https.git.savannah.gnu.org, and one or more of them seem to be outdated or not working right. As far as I know, Git itself provides very few tools to examine or control which host the fetching process uses; the best option 'git fetch' has is to select only IPv4 or IPv6 hosts (well, IP addresses).

(Quite reasonably, git fetch's verbosity settings are focused on the Git side of things, not on the network side of things. The network is supposed to just work, or at least fail in an obvious way.)

Fortunately we can take advantage of the simple Git HTTP protocol to directly query every server to see the state of their repository (assuming that they respond). Specifically, we want to use dumb client reference discovery to see the commit ID of one or more references (most often branch heads) on each server. To do this we'll need some way of forcing a HTTPS server name to resolve to a specific IP address, but curl has this feature in the form of its '--resolve' command line option.

(Curl has two ways to remap a HTTP server name; for using a specific IP address, --resolve is easier or at least more obvious than --connect-to.)

So what we want is something like this (assuming we care about the state of the main branch; you can pick another one):

host=https.git.savannah.gnu.org
url=https://$host/git/emacs.git/info/refs
ipv4=$(dig +short a $host.)
# Curl requires IPv6 addresses as
# '[...]'.
ipv6=$(dig +short aaaa $host. |
       sed -e 's/^/[/' -e 's/$/]/')
for i in $ipv4 $ipv6; do
  echo $i:
  curl -sS -L --resolve $host:443:$i $url |
    grep refs/heads/master
done

(I'm using 'dig +short' in this example as the most convenient general way to get the IPv4 and IPv6 addresses of the host, without anything else.)

At the moment, this says that all of the IP addresses are actually responding to Curl and one IPv6 address is outdated (ie, it has a different commit ID for refs/heads/master, and that commit ID is an old one). I will leave a nicer output format as an exercise to the reader; this is a quick hack that I'm writing down in case I ever need it again (and I hope not to).

(One improvement would be a script that you ran in a repository so it could look up the current head commit and only show you mirror hosts that had a different commit ID for their head.)

Actually doing anything with this information is also left as an exercise to the reader. As far as I know, Git doesn't let you not connect to one specific IP address, so you're left with more system level things like blocking connections to the errant mirror host. Right now, I'm just going to remember to use 'git fetch -4 savannah' when fetching from the official GNU Emacs repository (and hope that no IPv4 mirror host goes bad).

(If you're operating mirror hosts you can use this approach to monitor whether all of the hosts are sufficiently up to date and check for persistently out of date hosts. Or you may have a better monitoring method, for example based on internal mirroring data.)

A mistake I've made with the Apache IfModule directive

By: cks

Suppose that you (I) write an Apache configuration stanza to make some settings conditional on the module that they're from, so you can enable and disable the module without blowing up your web server configuration (or having to edit it). Your version looks like this:

<IfModule mod_qos>
  QS_LocRequestLimitMatch "^...$" 1000
  QS_SrvMaxConnPerIP 8 100
</IfModule>

Unfortunately, this stanza isn't doing what you (I) think it is, although it looks like it's correct. As covered in the <IfModule> documentation, the name you give IfModule is either a module identifier or a module file name ('the file name of the module, at the time it was compiled'). The 'mod_qos' I've used here turns out to be neither; the correct module file name for mod_qos is 'mod_qos.c' (or at least I think it is), while the module identifier is 'qos_module'.

(As covered in the documentation for LoadModule, you can get the module identifier by looking at the first argument to LoadModule for the particular module. Although I believe that people writing Apache modules can do it differently if they want to, the standard form seems to be <whatever>_module. I don't know if there's any good way to get the module file name other than guessing it's the conventional name of the module (or its .so file) plus '.c'.)

The effect of an <IfModule name> for a name that's neither a module identifier nor a module file name is that you've completely disabled that stanza, since the <IfModule> can never match an enabled module. You might wonder how you can make this mistake, and in my case it's simple. If you started out with an unqualified set of (module) directives and you're adding the <IfModule> with the intention of then (temporarily) disabling the module, well, the module configuration will be ignored after your 'a2dismod' and Apache restart just as if you'd gotten it right. You'll only discover the mistake when you try to enable the module later (or copy the configuration stanza to another web server entirely), and that might be years later.

(In our case we 'temporarily' disabled mod_qos on our web server in December of 2022 and then never re-enabled it because our web server stopped getting obviously overloaded.)

An unusual way for your DHCP server to run out of dynamic IPs

By: cks

Today I shared a brief war story on the Fediverse:

Today's new and exciting failure mode for a DHCP server handing out leases to dynamic clients: have something on your network that answers pings for absolutely every IP address (yes, it was broken). The ISC DHCP server pings what it thinks is a free IP before handing it out (to be sure), so if something answers all of those pings you have no 'free' IPs on your network and no one gets a dynamic IP.

(Technically this was like a day or two ago.)

The direct symptom of this in your ISC DHCP server logs is some log lines that look like this:

dhcpd[1656384]: Reclaiming abandoned lease 172.17.101.132.
[...]
dhcpd[1656384]: ICMP Echo reply while lease 172.17.101.132 valid.
dhcpd[1656384]: Abandoning IP address 172.17.101.132: pinged before offer

As ISC dhcpd documents (for example in dhcpd.conf's discussion of the 'ping-check' statement), by default dhcpd will ping an IP it's about to dynamically allocate to make sure it's unused. If something answers, dhcpd more or less gives up on the IP address (this doesn't happen for statically assigned IPs, at least according to the dhcpd.conf manual page). The consequence of this is that if you have such a 'screaming' machine, one that's answering ICMP pings for all IP addresses, dhcpd will conclude that your dynamic IP address pool is entirely exhausted and no dynamic client will be able to lease a new IP. For extra fun, apparently some clients will not accept a DHCP IP if there seems to be something else using it.

(I'm not sure what happens when clients are renewing leases.)

Such a screaming machine is obviously broken in some way and you need it off your network, but unless you get lucky, tracking it down may be hard. We were lucky that the machine was using its real MAC to answer all of the pings (which showed up all over the DHCP server's ARP table, among other places) and that MAC was registered with some useful and accurate additional information. Without that we would have been reduced to tracing through switch ARP tables (for switches smart enough to report that) and eventually unplugging sections of this particular network, which would have been pretty disruptive.

This particular network is port isolated, but that doesn't help here. Our DHCP server has to be able to reach the entire network and the entire network reach it, so its ARP requests flood through the network and anyone can answer them (and then its ICMP ping will be routed through the switch fabric to whatever port answered).

BMCs and a surprising USB network device on your server

By: cks

Suppose, not entirely hypothetically, that you're installing a server with two network ports and during the (Linux) installation, a third network device shows up with a funny name like 'enp1s0f4u1u2c2', which is a USB Ethernet device (despite you not having any such thing plugged in to the server's USB ports). To your further surprise, your server installer can even lease a DHCP IP on this interface, say "169.254.3.1". Congratulations, your server has a BMC, and this BMC probably speaks Redfish, which is sort of the modern, cloud influenced version of IPMI.

One of the things you'd like to do on a server with a BMC is have the server (the 'host') talk directly to the BMC, for example to get sensor information that only the BMC has or to configure the BMC. In the world of IPMI, people had to put together special methods to talk to the BMC, which required kernel drivers, extracting information from SMBIOS, and so on. This isn't the greatest, and is also not at all like how you talk to the management agent in cloud virtual machines, where you generally talk to the management agent by making HTTP requests to a special IP address. IPMI is in part a network protocol, but talking to a BMC using IPMI over the network is completely different from talking to it from the host server.

The IPMI protocol is an essentially custom UDP based thing, which made sense at the time. Redfish instead uses a HTTP REST based approach, partly because by the time Redfish was started, it was obvious that HTTP had become basically the universal protocol (and it was already in use for similar management purposes in cloud environments). So the natural way for a host server to talk to its Redfish based BMC is over some sort of network connection, instead of through some special out of band mechanism the way IPMI does. However, this requires a network interface that's directly connected to the BMC (and nothing else).

You could in theory wire up some sort of semi-virtual PCIe Ethernet device that was connected to the BMC on the other side. But that's complicated. Most BMCs support 'KVM over IP', and as part of that they need to provide virtual keyboard and mouse input, which these days is done by having the BMC present a (virtual) USB keyboard and mouse to the host. Many BMCs can also present USB storage media to the host, for install media. If a BMC is already presenting a bunch of virtual USB devices to the host, the obvious way to provide a network interface to the host for Redfish is through a virtual USB Ethernet device.

(I think all of these virtual USB devices are often presented on a virtual USB hub, and maybe even a virtual PCIe USB controller to go along with the virtual PCIe graphics card and maybe PCIe bridge. There's a lot of funny business that goes on to connect a BMC to the host system, never mind issues like how BMCs may control the host power.)

Since the host needs to have an IP address on this virtual USB Ethernet device to talk to the BMC at the other end, the BMC has a little DHCP server as well as its HTTP server. This internal HTTP server may or may not be the same as the BMC's regular management HTTP server. On some of our servers, this internal BMC HTTP server only answers Redfish requests and doesn't provide the normal BMC web interface that you can get on the BMC's management network interface (which also supports Redfish requests, of course).

(I think this is a sensible security decision on the BMC's part.)

Notes on pulling from multiple upstream Git mirrors

By: cks

It started with a discovery about my access to the official Emacs repository:

This is my face when my home desktop appears to be persistently talking to one instance of https.git.savannah.gnu.org that is many days out of date and out of sync on the GNU Emacs git repository. Yes, I know, volunteer organization, but how do you even troubleshoot that? At this rate I'm going to have to switch to the Github mirror even though the thought makes me spit reflexively.

(By 'how do you even troubleshoot that', I meant how might I figure out which mirror is out of date and report it. That DNS name has eight IPv4 addresses and eight IPv6 ones, although you can at least restrict Git to either IPv4 or IPv6.)

This led to me wishing for a way to conveniently pull the same branch from two different upstreams. To be specific, the experience I would like looks like this:

$ git status
On branch emacs-31
Your branch is up to date with 'origin/emacs-31'.
[...]
$ git pull
[pulls and updates my emacs-31 local checkout
from some reliable mirror]

$ git pull savannah
[pulls from the official repository and also
updates my emacs-31 local checkout]

As was pointed out to me by several people (once I read the git-pull manual page), you can get almost this experience with plain remotes, but on your non-default remote you have to remember to use a special form, 'git pull savannah emacs-31'. As far as I can see there's no way to tell 'git pull' to do this automatically, since 'git pull' goes from your local branch to the remote (and you can only have one remote).

(At this point you could make a git alias for this specific operation, perhaps called 'git alt-pull'.)

You can also merely fetch from the non-default upstream and then manually trigger the same nominal merge that 'git pull' would:

$ git fetch savannah
$ git merge --ff-only savannah/emacs-31

I also had a pseudo-clever idea that almost certainly won't work, as I can see better now that I've read a bit more about 'git pull':

I could manually edit .git/config so that both the savannah and github remotes updated the same local 'refs/remotes/origin/*' ref(s), but I suspect that this would go badly and also not necessarily ripple through to updating the on-disk state when I did a 'git pull' from the one that isn't listed as 'remote =' for the emacs-31 branch.

Manually switching the remote of the emacs-31 branch back and forth will do that but, well, annoyance.

Given how 'git pull' works, I believe this would give me the same 'you asked to pull ... but didn't specify a branch' error as 'git pull savannah' does in a standard configuration. It's possible that 'git fetch savannah' followed by 'git merge --ff-only' would work (if the shared remote HEAD was updated by my fetch), but that's already two commands and not much different than other options (at the cost of possibly confusing Git).

Manually switching the remote of my local branch back and forth can be done on the command line with git's general '-c <name>=<value>' setting for setting a configuration parameter:

$ git -c branch.emacs-31.remote=savannah status
On branch emacs-31
Your branch is ahead of 'savannah/emacs-31' by 17 commits.

Once again I could make a cover script that set this for all Git commands, or I think I could do it for specific commands through Git aliases (which I think would make it easier to pass command line arguments through compared to a 'git alt-pull' alias).

(The real answer is that I'm likely to switch more or less permanently to one of the mirrors and give up trying to directly fetch from savannah.gnu.org, which is apparently very overloaded and not all that healthy. One reason I cloned directly from savannah is that I had the impression that the mirrors could lag significantly behind, but on a spot check today the Github one was pretty up to date, with commits only an hour or two old. But at least going through this exercise has left me a bit more educated about some Git stuff.)

Our mixed building network wiring and its consequences

By: cks

In a comment on my entry about how sometimes it's the network that's the problem, I was asked what networking workstations typically have on our networks (and if we still use 100 MBit for low speed things). The simple answer is that it's somewhat mixed but mostly 1G Ethernet, often with 1G uplinks from the switches that the office network jacks are connected to. But there's some stories involved.

As a university department, we have been in our current buildings for what is mostly multiple decades, and in some cases since the early 1980s (cf). In the oldest building (with our old machine room), the department's presence predates twisted pair network wiring entirely (there is still disconnected thicknet wiring around our offices); our other most used building predates Cat-6 cabling. As a result, all of this older space is wired with Cat-5 twisted pair and is only good for 1G Ethernet.

We started out using all of this wiring at 100 MBit (including our connection to the university backbone), and even as time went by we had a mixture of 100 MBit and 1G switches and connections. Due to cost reasons and sometimes wiring density problems, we only slowly migrated people's office network jacks from 100 MBit to 1G as we progressively replaced switches in our multi-switch setup. However, we did eventually move everything to 1G, including in our machine rooms (partly because our old 100 MBit switches were getting unreliable). I think most of our '1G' switches in wiring closets and so on are purely 1G, with no 10G uplink, so wired workstations on those 1G network jacks are sharing the uplink bandwidth.

Areas in buildings do get renovated from time to time, including parts of the department's space. For years, these renovations have typically included running new Cat-6 (or Cat-6A) wiring that is properly certified and tested for 10G-T (it's hard to significantly renovate space without needing to tear out the old wiring and then put in replacement wiring). Sometimes the people who are funding the space renovation really do want their desktops to have 10G connections available, and in that case the renovation funding will also include a collection of 10G-T switches for the other end of the office network jacks to plug in to (however many switches needed for however many true 10G office ports, which is often all of them). Otherwise, while we have 10G-T capable wiring in the walls, we only plug our end of the wires (in wiring closets and machine rooms) into 1G switches and everyone gets 1G.

(Newly built out space in new buildings is all wired with Cat-6A, of course, and we try to get funding for 10G-T switches to go with it. There's been some of this over the recent enough past, as the university does acquire a certain amount of new space over time.)

All of this is the theoretical answer. The practical answer is that a lot of people are using laptops (possibly with docking stations) or desktops with wireless built in, and many people default to using the wireless connection even if there's a network jack or three in their office. Typically, the only people who use actual wired connections are those that need their machines on one of our special research group private networks, instead of on our general 'random desktops and laptops' network. In some areas of the department, there's almost no use of wired network jacks because almost everyone is using laptops and the wireless.

(This makes wireless into critical network infrastructure.)

Here in 2026, I wouldn't try to use 100 MBit for anything, even things that don't need 1G, because I wouldn't trust it. In theory your 1G network interfaces and switches still support it (and maybe your 10G ones too, but don't count on it). In practice, it's been a fairly long time since 100 MBit hardware was in common use. Any switches or the like that you still have lying around are old, and these code paths and physical capabilities in network ports aren't likely to be heavily tested. You want to be running your networks in common and well supported setups, and today that means 1G is the minimum speed you should target.

(And if you have cabling that seems to only support 100 MBit, it's probably broken. Network cables do go bad over time.)

Using Linux tc to limit the outgoing bandwidth of a web server

By: cks

Suppose, not hypothetically, that you have a web server that's using as much of your server's bandwidth as it can get and you would like it to use less bandwidth than that, so that you can get a word in edgewise (for backups, for example) or just because you don't feel like donating 1/10th of your outgoing bandwidth in apparent perpetuity to people who should be building local caches. There are various ways you might do this, for example using FreeBSD pf on your perimeter firewall, but the lowest impact and risk option is to do it on the (Linux) web server itself with tc(8), the Linux traffic control system. Conveniently I've already done a tiny bit with tc to fight bufferbloat latency.

There are probably a variety of ways to do this in tc(8), but what I'm using right now is mostly pulled from the Arch wiki. It goes like this:

  1. We have to switch our device, which is eno1 for me, to using Hierarchy Token Bucket as its top level qdisc (queueing discipline), and in the process set a default class that otherwise unclassified packets will be assigned to.

    tc qdisc del dev eno1 root
    tc qdisc add dev eno1 root handle 1: htb default 30 r2q 1000
    

    Following the Arch example, the default class is 1:30 (the root handle of '1' combined with 'default 30'). Because I'm working with high bandwidth, I need to change r2q to make tc happy.

  2. Set up a child class to limit bandwidth, here with a limit of more or less half of a 1G Ethernet. Because we're only doing one level bandwidth limiting (ie, we're not sub-dividing it), this can go right under the root parent ('1:').

    tc class add dev eno1 parent 1: classid 1:10 htb rate 400mbit ceil 500mbit prio 10
    

    Pick the bandwidth numbers to taste depending on your irritation (and the interface speed of your server, your outgoing bandwidth in general, and so on).

    The classid is somewhat arbitrary but as the Arch Wiki example shows, there can be a use for setting a numbering hierarchy if you have multiple levels of parent and child classes. In the Arch example, the top level bandwidth limited class is '1:1', then it has child classes '1:10', '1:20', and '1:30' (the default class for traffic).

  3. Following the Arch Wiki example, we add a fair queueing qdisc below our bandwidth limit, so traffic flows that fall into this bandwidth limit are (hopefully) a bit better handled. I think this means that one HTTPS reply to a requester on a fast link won't starve all of the others.

    tc qdisc add dev eno1 parent 1:10 handle 10: sfq perturb 10
    

  4. Filter outgoing HTTPS traffic into our bandwidth limiting class:

    tc filter add dev eno1 protocol ip parent 1: prio 1 u32 match ip sport 443 0xffff flowid 1:10
    

    This is a u32 match, and it requires some decoding to understand. A u32 match like this is fundamentally 'match <value>/<mask> at <offset>', and if we dump the raw form of this with 'tc filter show dev eno1', we'll get "match 01bb0000/ffff0000 at 20". The 'ip sport 443' format is simply tc-u32's friendlier way of encoding that (well, technically 'ip sport 443 0xffff', since the 0xffff is a load-bearing part of the shorthand). The tc-u32(8) manual page has various cautions about this matching, so I think if you really care you want to use iptables (or nftables) rules to set a firewall mark and then match on that with tc-fw(8).

  5. Create our default catch-all class that effectively has no bandwidth limit, and as with our bandwidth limited class, also attach a tc-sfq(8) qdisc below it.

    tc class add dev eno1 parent 1: classid 1:30 htb rate 1gbit ceil 10gbit prio 1
    tc qdisc add dev eno1 parent 1:30 handle 30: sfq perturb 10
    

    We have two choices for 'prio'. The first option is that we can set this to 'prio 10', the same as our bandwidth limited class; in that case, I believe non-limited traffic will share bandwidth with the bandwidth limited HTTPS traffic, basically getting what's left over. Alternately, we can decide that we want non-limited traffic to have priority over bandwidth limited traffic, in which case we want a lower priority so its packets are sent first. That's what we're doing here.

    We need a classful qdisc, but we'd like one that will automatically use all of the bandwidth available to it. I've used tc-htb(8) here because it's what I was already using, but there are probably better options. The bandwidth I put here is simply 'really big', starting with the 1G interface rate.

This appears to work but the tc-htb(8) rate numbers may not correspond exactly to the wire bandwidth you (I) see. For our purposes, this is good enough; we're not trying to limit things to an exact MBytes/sec value, just something that's in the right ballpark.

Now that I've read the tc-u32(8) manual page, I'm pretty certain I wouldn't want to use it for anything significant where I cared deeply about what traffic was getting sorted into what class. Iptables or nftables is going to be much better at matching network traffic, especially in unusual and weird situations, and you can use that in traffic control rules through tc-fw(8). Here it doesn't matter too much if some unusual HTTPS traffic 'leaks' outside of my tc filter, as long as there's not too much of it (plus the whole server has an intrinsic bandwidth limit).

PS: For our purposes it would be nice to do something sophisticated to aggregate all HTTPS requests from a single IP address together in a single 'flow' for tc-sfq(8). In theory this is possible with tc-flow(8) but in practice I can't figure out a command line that works right (despite consulting eg tc-sfb(8)'s example). I'm sure the documentation makes sense to people who have a deep understanding of Linux traffic control, but I'm not such a person.

Sometimes it actually is the network: a war story

By: cks

We've recently been having mysterious problems getting some of our backups to perform well. Also I recently wrote about how we'd wound up with a web server that frequently saturated its outgoing 1G interface with traffic (and it was a feature that it didn't have a faster network link). These two things turn out to not be as unrelated as we'd like, and there's a story or two there.

The problematic backups aren't done through our usual Amanda-based backup system; instead, a particular central machine pulls /var/log and other relevant files from FreeBSD and OpenBSD hosts (either through rsync or direct SSH access) and then writes them over NFS to one of our NFS fileservers (and then they get backed up by the Amanda backups). The problems manifested as terrible NFS performance; the central machine's load average would go to 50 or 70, everything NFS related on it would be slow, and so on. While we like NFS in general we've had quite a share of problems with it (cf, also), so we immediately assumed that this was yet another instance of an NFS problem. First we reduced the IO load from rsyncs in various ways (eg), then we spent quite a while digging at various NFS metrics, both from our metrics system and from the live system during a problem. Nothing really seemed to be the problem; the fileservers had perfectly good performance, the master machine had perfectly terrible actual NFS performance, and especially it didn't seem to be able to write the backup data very fast, in the range of a few MBytes a second.

Since I was pretty sure that the central machine could write to its local disk acceptably fast, I considered switching to a more complex backup scheme where we first rsync'd things to the local disk, packed up this directory tree into a tar archive, then scp'd it to the relevant fileserver, thereby bypassing any NFS write issues. To check that this would work acceptably fast, I started by scp'ing a test file from the central machine to the fileserver. To my surprise, the scp ran at only about 2 MBytes/s. More testing showed that scp's from this central machine to anywhere ran at 2 MBytes/sec at most, which did rather explain the NFS write problems (much like the lack of a CPU explains a server's failure to power on). At this point a penny dropped in my mind.

We're a university department, which means that we don't necessarily have the newest, shiniest stuff around and we keep things in service for a long time. One of the things we've kept in service is basic 1G switches, because quite a lot of servers don't need more than 1G (and as mentioned, sometimes it's a feature that a server only has 1G). But when I say 'basic 1G switches' I mean switches that are so old that all of their ports are 1G, including the port we use for 'uplinking' them into our overall 10G switch fabric (in contrast to modern 1G switches, many of which have one or two SFP+ ports that can run at 10G, or even a 10G-T port or two). This is fine for our normal 1G servers, which don't generate or receive much traffic even in aggregate, but it breaks down badly the moment you put a high volume 1G server on such a switch. For example, our new high volume web server, which was not only saturating its own outgoing 1G interface but was also saturating the 1G uplink from the switch it was connected to, a switch which unfortunately also had this critical central machine connected to it.

There's a programmer saying that if you think you have a compiler bug, you don't, you have a regular bug that you haven't spotted yet. This saying is almost always true, and it generalizes to other areas, like kernel bugs (also) or hardware problems, or networking problems. At our scale, modern networks are reliable, so if we don't have anything obvious wrong (for example, our monitoring system hasn't alerted us that some machine is unexpected at 100 MBit/sec), our network is almost certainly working fine. So for days it didn't occur to us to actually check. Of course the network was working fine, the network is always working fine. It obviously had to be NFS, especially since NFS has been flaky for us in the past. Except that sometimes it is the network and there's even a good explanation for it that becomes obvious once you realize it's the network.

This whole experience has given me some things to think about. On the one hand, years ago I wrote about how an obvious problem isn't necessarily obvious. There are a ton of things that can be wrong and you have to winnow through them somehow. On the other hand, I'm pretty sure that if we'd engaged in systematic troubleshooting from the ground up, we'd have found this pretty early on. For example, the USE method would have had us look at usage, saturation, and critically 'errors', which might well have caused us to look at TCP retransmits on the central machine (which were decidedly high).

Being systematic about troubleshooting is generally a good thing, but at the same time it's tedious. If the problem had really been a NFS problem (as it has been in the past) and I'd followed the USE method from the ground up, I'd have spent a chunk of time verifying that yes, the network was performing fine on both machines (along with the other things I looked at it, like NFS server metrics). Possibly what I should try to do is start out with likely guesses and then when they come up dry (eg, there are no obvious reasons for a NFS performance problem) and I'm getting frustrated, fall back to the USE method or something similar, even though it's possibly tedious.

What buffer size (OpenSSH) ssh seems to use for streaming output

By: cks

Suppose that you're generating and transferring a file over ssh, for example to create a tar archive of something on a remote server and save it locally:

ssh [...] rem-server 'cd /var/log && tar -cf - .' >/backup/file.tar

If you're experiencing IO problems in this backup process, an interesting question is what buffer size ssh uses for its writes, perhaps because you'd like to make a few large writes to disk (for example, at the natural 128 KByte block size for your ZFS fileservers, even when you're writing over NFS) instead of a bunch of smaller ones.

The ssh manual page doesn't document anything about this, and there's no options to control it in either ssh or ssh_config(5). On Linux, running over TCP from a remote machine, the answer appears to be that ssh will normally do 32 KByte writes (or 64 KByte writes under some circumstances, possibly if the network is fast enough). It's possible ssh will write smaller buffers if the remote command generating the output can't keep up at full network bandwidth, and in general I suspect that there are a lot of things that can change this.

If this is important to you (and I'm not convinced it's important to us), you need to re-buffer the output from ssh in some way. The traditional way is to use dd, but you need to pick the right options to have dd reblock its input. There are other programs floating around but I don't know if any of them are standard.

(Since I looked it up, there's at least the Debian buffer package and mbuffer. There are probably others out there as well that I can't dig up in casual Internet searches.)

Finding out this information requires some way to trace ssh's system call activity. On Linux this is most easily done with 'strace', and you can narrow down all of the system calls that ssh does with 'strace -e trace-fd=5 ...'. You might think that you want to trace file descriptor 1 (standard output), but in fact current versions of OpenSSH ssh rewires its standard output on to file descriptor 5 and make file descriptor 1 point to /dev/null (which can be very confusing when you first encounter it).

(This is one of those things that I look into, don't find much, and then want to write down my negative results anyway for future use.)

Go maps, hashes of map keys, and pointers: a little surprise

By: cks

Go's maps are famously implemented as hash tables, which is the only reasonable choice. The implementation has gotten somewhat more complicated since I looked at how maps store their values and keys due to the move to swiss tables, and these days you find the comments about how they work in internal/runtime/maps/map.go, but the core is still the same. Recently, a documentation commit landed in the Go development tree that opened my eyes to a bit of subtle complexity I hadn't considered before in Go's map implementation.

One of the things about hash tables is that they hash the value of keys down to some fixed size value in order to do operations more efficiently; in Go's current swiss tables, this is a 64-bit hash. Critically, the hash value of a key must be constant (which can be an issue in languages like Python that let you define a hash function at user level). You also want the actual value of keys to (only) compare equal when they are equal, which can also be a challenge in a language with user defined comparison functions, or just if you're dealing with NaNs.

Go has a quite broad definition of what's allowed as map keys; you can use any type that has == and != comparison operators defined. This includes pointers (which are directly comparable), arrays of pointers, and structs containing pointers (under the rule that a struct is comparable if all its fields are). However, Go pointers aren't guaranteed to be constant values, and today growing and shrinking a goroutine's stack will change some pointer values. This is a potential problem if you're hashing the current integer value of a pointer as part of a Go map key hash; you need that hash value to stay constant.

The documentation commit explains how Go deals with this today, primarily in its comment in map.go. When a Go value is stored as a map key, the Go compiler marks that value as 'escaping', which means that the value will be allocated in the heap instead of on the stack (along with anything it points to). Currently things in the heap never move, so once a key value is heap allocated, any pointers involved have a constant value and the key's hash value will never change.

As the comment notes, this is only done for keys that are getting stored in the map. Keys used for lookup or for delete will never be stored and so don't need to be specifically heap allocated. As the comment also notes:

If we are looking up a pointer which points to the stack, the hash value is ~irrelevant, as the key is guaranteed to not be in the map [...].

(This includes pointers in structs and so on.)

This map key hash stability requirement is a bit of a subtle constraint on any future Go garbage collector that works through copying values around (eg, also). Probably the simplest way to deal with it would be to mark heap pointers involved in map keys and then never copy or otherwise move them. Possibly you could do this on the fly during the garbage collection scanning process, since you need to trace through maps in general to mark their keys and values as used.

(Until I stumbled over this commit message and read into it more, I'd never thought about how the hash table stability requirement might clash with any sort of moving garbage collection mechanism.)

Discovering rsync's -W option and our use for it

By: cks

Suppose, not hypothetically, that you use rsync to push an encrypted backup file from the machine it's created on to a fileserver, where it will be backed up by your regular backup system. Because this encrypted backup file is backed up every day, you only need one copy of it on the filesystem, so you use and reuse a fixed name for the file. In other words, we're using rsync somewhat as if it was scp, but with better control over what remote files can be written and (not) read.

When you push (or pull) a file over rsync, rsync normally attempts to optimize what gets transferred by looking for common blocks in the file (how big a 'block' is depends on the file size, or you can fix it with the '--block-size' option, as covered in rsync(1)). This is a nice potential bandwidth saving, but it creates CPU and IO load on both ends as each of them checks their version of the file. In our specific case, we know that there aren't going to be common blocks; since the whole file is encrypted, it's basically random noise. Recently this backup process had some IO load problems, and today in the process of working on this I discovered rsync's '-W' option (also known as '--whole-file'). As the manual page explains, this 'disables rsync's delta-transfer algorithm'; in other words, it stops looking for common pieces between the two versions of the file. Rsync simply sends the whole file (and the receiver writes the whole file).

Since we know that today's encrypted file has no blocks in common with yesterday's encrypted file (well, it had better not if the encryption is working right), '-W' is exactly what we want to stop the receiving rsync daemon from doing unnecessary work (specifically, unnecessary IO). Effectively it turns the 'file copy' part of rsync into scp (although not literally; rsync will normally write the new version of the file to a temporary file and then replace the old version). Now that I know about -W, I'm going to be looking at some of our other uses of rsync to see if we might want to use it more widely.

(For example, we use rsync to back up /var/log from some FreeBSD hosts, and I'm pretty sure that's a good candidate for -W too.)

If you're using '-W', you want to avoid using '--checksum' and instead rely on the default 'quick check' of file size and modification time. This is because using --checksum requires rsync to read and checksum the whole file before the transfer starts (which is something that the manual page warns you about).

One lesson I've taken from today's experience is that when I use rsync, I should think about what I want to optimize (and what can actually be optimized). Rsync's default behavior is to optimize transfer bandwidth, but sometimes you have enough transfer bandwidth and you want to optimize for lower IO, lower CPU, or both (which is sort of our case for this encrypted backup file, with the extra issue that we know rsync can't reduce the transfer size). Alternately, sometimes you really want to squeeze the bandwidth and maybe '-S' and '-z' (and perhaps others) are what you want, even though you'll do more work on both ends.

(It's possible that rsync already has a clever encoding for runs of zero bytes and so '-S' doesn't save you any transfer bandwidth. I haven't tested.)

(This elaborates on a Fediverse post of mine.)

I'm only interested in "native" installation systems

By: cks

There are a variety of ways to automatically or semi-automatically install systems, especially Linux systems and especially over the network. People have built a whole raft of them over the years, often with relatively impressive capabilities. A number of them have significant levels of automation and control, including for things we might want like sophisticated automatic disk setup. Despite that, we're interested in approximately none of them. As a practical matter, we only want to use the standard, native install systems for whatever we're running, which in this case is the Ubuntu server installer. There are two reasons for this.

The obvious reason is that only native installers and install methods are officially supported by the Linux distribution (or whatever other free Unix we're using), and they're also the ones that are most used in practice. The native installers aren't always perfect, but other people work (often quite a lot) on making them work well, we can find plenty of people using them, and so on. If we use another installer, at the very least we're without support from the distribution. We're also likely to be off the beaten path, which means fewer people running into issues before us, fewer bug reports being filed and fixed, and so on.

The less obvious reason is that we do a certain amount of direct manual installation of servers for Ubuntu, FreeBSD, OpenBSD, and so on. The easiest way to do this is to use whatever native installer the distribution ships on their install ISO images. In the case of Ubuntu we customize this a bit, but none of our customizations are essential; we can and have installed systems straight from the main Ubuntu ISO images (and then possibly imported our customizations afterward). We could still do this even if we normally used a third party install system, but less would carry over between the two environments and we'd want to learn and stay familiar with both of them.

These issues aren't absolute blocks on using a third party installer system; we could deal with both of them. But we'd want to be getting something important from such a system, something that was enough of a gain to be worth the extra costs. I can think of situations which would be worth it but they haven't come up so far (at our modest scale and with only Ubuntu really in the picture for semi-automated installs, network installs, and so on).

(This is related to why it took us a relatively long time to build up a network install system, cf. There are specialized automated network install systems for Linux, but they run into this third party issue. Our current network install system for Ubuntu uses our Ubuntu ISO images that we also use for local installs, and that's a feature even though it creates complications.)

Our long path from IPMI remote installs to network installs

By: cks

Once upon what's now a long while ago, we had a bunch of SunFire X2100 and X2200 1U servers (for example, they were used in our first generation ZFS fileserver). One of the reasons that I loved these servers back in the days was that for free, they came with a full IPMI/BMC setup that supported both KVM over IP and virtual media. I installed any number of servers this way from the comfort of my office, using the KVM over IP as if I was at the console (because I effectively was) and the IPMI virtual media to feed a local ISO image to the server I was installing. In time those servers went the way of all servers (which is to say, into the e-waste dumpster, although it took a while) and none of our later servers gave me that '(re)install from the office' experience with their limited BMCs. For years, we lived with doing in person physical installs and reinstalls of our servers (well, I lived with that, my co-workers don't care as much).

Recently we wound up with a bunch of servers in another building and a need to reinstall them all with Ubuntu 26.04. This pushed me into learning about UEFI network boot, working out how to network boot our customized Ubuntu ISO image, and then the realization that if we were reinstalling an existing, running server we could use a somewhat simpler kexec-based method. This has given us a reinstall experience (and sometimes an install experience) that looks a fair bit like the SunFire X2100 BMC based experience, and can be done from comfort of our office (instead of a noisy machine room a block or two away). My co-workers like it enough that we're talking about using our new network reinstall system even for servers in our main machine room.

On the one hand, it's nice that we can now have this experience with basically any system (it's nicer if it supports network booting, but reinstalls work without that). On the other hand, the experience is a lot more fragile than the BMC-based experience. Our network installs only works because the Ubuntu server installer supports access over SSH, it requires either a running system on the server or a whole collection of network booting infrastructure (including that it be enabled on the server). If the installer blows up (as it does too often in 26.04), you can't necessarily get access to the system or force a power cycle, and you may be stuck needing to go visit the server in person.

(The most practical version of the network install experience also relies on the server having enough RAM to download the ISO image. This seems to require 8 GB of RAM in common situations in my testing, which is far more than our SunFire X2100s normally had.)

This change from hardware support on limited machines to general software that more or less replicates the earlier experience (but imperfectly) feels like a pattern that's happened repeatedly. One advantage of the software version is that it generalizes, for example to virtual machines.

(Virtual machines can have a BMC-like experience, but it's all with a bespoke environment that's specific to that VM system. Network installs are generic across physical servers, virtual machines, and so on.)

PS: Of course the SunFire BMC used a Java applet, which was somewhat painful (cf). Modern BMCs thankfully use straightforward web stuff.

As expected, using kexec to switch to a new Ubuntu kernel works

By: cks

For a long time, I ignored kexec both for my own personal machines and at work. I knew it existed but I never attempted to use it. That changed recently when I realized we could use kexec to start a network (re)install environment without needing our servers to actually have network booting enabled (which they mostly don't currently, for historical reasons). This has led me to do some additional experimentation with kexec.

The most recent experiment was updating to a new Ubuntu kernel and then using 'kexec' to switch to it instead of 'reboot'. On Ubuntu, this can be pretty simple due to some short symlinks, more or less:

kexec -l /boot/vmlinuz --initrd /boot/initrd.img --append "$(cat /proc/cmdline)"
systemctl kexec

But I'm not sure we're going to actually use this in anything except very unusual circumstances.

The good side of using kexec instead of reboot is that you don't have to sit through your BIOS slowly fiddling around with hardware and then GRUB doing GRUB things. This gets your system back on the air anywhere between 'somewhat faster' and 'much faster', depending on how slow your BIOS firmware is (and also your GRUB timeout). Some of our servers have firmware that is pretty close to 'agonizingly slow', which is part of why we often have network booting disabled on them (this often skips a slow BIOS setup process for network cards, even if network booting was low on the boot priority).

But this good side is also the issue with not using reboot, because if you don't use reboot, you haven't actually tested and verified a cold boot situation. You don't know that your new GRUB boot entry for your new kernel works, and you don't know if your firmware is going to discover some hardware problem at an inconvenient time later on, instead of now during a planned downtime when you're available to deal with the situation. Of course probably the GRUB entry works and probably there's not going to be any surprises from the firmware (and maybe you can inspect the GRUB entry yourself to be sure).

For us, verifying the cold boot behavior (and having a simpler and less error prone 'switch to new kernel' process) is almost always going to be more important than a fast reboot. Now that we know about it, we may kexec into new kernels in exceptional circumstances, but I doubt we're ever going to do it routinely.

There's a plague of Googlebot impersonation going on (in June 2026)

By: cks

A while back I wrote about how claiming to be Googlebot is now a bad idea, where I noted that there were (reports of) malicious crawlers out there impersonating Googlebot and other legitimate big crawlers and at the same time, Google and other crawler operators published the IP address ranges their crawlers used. You could put these two together to block these impersonators:

Anything claiming to be Googlebot that is not from a listed Google IP is extremely suspicious and in this day and age of increasing anti-crawler defenses, blocking all 'Googlebot' activity that isn't from one of their listed IP ranges is an obvious thing to do.

After I wrote that entry, I quietly went and added support for blocking crawler impersonators to DWiki, the wiki-engine that powers Wandering Thoughts, and set it up for a few big crawlers with published IP address ranges. When I did this, I didn't expect to block very much, and for months that was indeed what happened; I'd get a few attempts once in a while. Then, this June, the floodgates opened.

For weeks, I've been seeing hundreds of requests a day claiming to be Googlebot (on a few days, thousands of requests). The requests come from a variety of IP addresses at a variety of providers, which I think are mostly or entirely cloud and hosting providers. The top sources by ASN are a rogue's gallery of places that I was already having problems with, like HostRoyale, M247, Latitude.sh, and Web2Objects. But AWS is in the collection as well, and there are a lot of other relatively mainstream providers. Many IPs seem to make only a few requests as Googlebot, and at least some of them immediately retry with another User-Agent value (which also generally doesn't work).

My guess is that this isn't a bunch of different abusive crawlers who've all spontaneously decided to try forging Googlebot to see if it gets them anywhere. Instead, I suspect that this is a large scale campaign by a single abusive crawler, run by people who can afford to obtain a lot of servers at a lot of different hosting providers (or who are prepared to commit various sorts of criminal fraud on a large scale). Ironically, if they'd picked a different tactic, I might not have noticed them among the background radiation of crawl attempts. Forging Googlebot and other known big crawlers is generally sufficiently rare that I actually bother looking at my logs to see it happening.

(I also suspect I'm not the only website this is happening to.)

PS: It feels somewhat ironic that this is happening at the same time as me wondering if I should allow Googlebot at all.

Go interfaces, reflection, and binary size

By: cks

Recently an interesting series of commits landed in Prometheus with the goal of reducing the size of the Prometheus binary by allowing the Go linker to remove more unused code (something it's quite good at in general, although the linker is also deliberately limited in this). The commit with the message that's most informative about what is going on and why is discovery/gce: keep [Google Cloud] Compute SD client from defeating dead-code elimination, and you can read the full details in it. The short version is that if you're using certain sorts of reflection anywhere in your program, the Go linker won't remove exported (public) methods of any concrete type that's reachable through an interface. It doesn't matter how narrow the interface is (it can be the famous and minimal fmt.Stringer); the moment you combine reflection and a concrete type in an interface, the Go linker more or less stops throwing out unused functions and methods. Well, sort of, as the commit explains.

Unlike the standard Go toolchain not doing dead code elimination for package level variables with constant values, this isn't merely the linker deciding it's too much work to do this dead code elimination optimization. Instead it's at least partly a correctness issue. The problem for the Go linker is that reflect allows you to reach through any retained interface value to use any and all exported methods on the underlying concrete type of the value (and any types it contains), using things like Value.MethodByName() and Value.Call(). This makes it hard or impossible for the Go linker to know which exported methods are really dead and can never be reached at runtime.

(This has to apply to concrete types contained in top level concrete types because reflect can reach through structs, channels, maps, arrays, and so on to retrieve underlying types and values, and thus methods on those types.)

The current Go linker is actually doing more work and eliminating more dead code than the documentation requires it to. The documentation for Value.MethodByName() and friends say that they apply to all exported methods (possibly only of a given name), but apparently the linker will skip this for types that are never directly or indirectly boxed into an interface, because such types aren't reachable through reflect. Since all reflect functions that create a Type or a Value take an any (ie, 'interface{}') as their argument, you can't go from a value of a concrete type to either without putting the concrete type in an interface and triggering this. What this means in practice in a program where there's any use of reflect (including in some sub-dependency off in a corner) is that if you put a 'big' type with a lot of direct and indirect exported methods into an interface, all of those methods and all of their dependencies will have to be retained in the binary (and increase its size, possibly a lot), even if you only use a tiny subset of them.

(I believe this includes innocent looking things like merely printing such a 'big' concrete struct, which you might do for debugging purposes or because it has a String() method that does useful stuff. And of course JSON serialization uses interface values; json.Marshall() takes an 'any' as an argument, so there's your interface. While the json package uses reflect internally, it doesn't currently call any of the reflect methods that triggers this linker behavior.)

There are at least two ways around this, visible in the Compute service discovery commit and a similar Kubernetes commit. In the Kubernetes commit, a concrete top level Kubernetes struct was not retained in full in a Prometheus service discovery struct that would then be boxed into an interface; instead, only the methods on the Kubernetes struct that were actually needed were extracted and embedded into a new struct, so the Go linker only had to retained those methods and their code dependencies. In the more complex Compute commit, some processing had to be done dynamically using concrete types that had to be retained, so instead of putting the concrete types in a Prometheus struct (that would then be boxed as an interface inside the Prometheus code), the values of the concrete types were made inaccessible to reflect by putting them inside a function closure, and only the function closure was stored in the Prometheus struct.

One thing I take away from this is that one should avoid using the various reflect method-getting methods if at all possible, both in a program and especially in a package that you expect other people to use. If your package uses these internally, you're creating spooky action at a distance effects on the whole program (and you should probably mention this in your documentation).

PS: The Go linker's dead code elimination is (currently) discussed in general in a big comment in cmd/link/internal/ld/deadcode.go, which is worth reading for some details that I hadn't thought about until now, such as needing to retain all methods that might be reached through interfaces (which is necessary because you might wind up casting an interface value to another interface entirely, eg, also).

PPS: As mentioned in the Prometheus commits, one of the packages that uses reflect this way is go.yaml.in/yaml/v4. For the actual code and usage involved, see here and here, which seem like reasonably sensible uses to me, even if they have awkward consequences.

More emulation goodness, an Intel Itanium (IA-64) emulator that boots Windows!

The emulation space is going crazy, after my previous [post on Windows booting on DEC Alpha es40 emulator](/s/blog/Run_Windows_2000_for_Dec_Alpha_on_a_new_es40_fork.html), there is now another huge breakthrough in the emulation of other non-x86 CPU emulation. [Yufeng Gao](https://github.com/TheBrokenPipe) with help from [gdwnldsKSC](https://github.com/gdwnldsKSC) (the man behind the updated es40-fork) has released version 0.1 of his [Intel Itanium (IA-64) emulator](https://thebrokenpipe.com/ia64/zx2000/) that boots the Itanium version of Windows Server 2003 and Windows XP 64-bit. No OpenVMS or HP-UX yet and Linux / BSD also don't boot. But Windows is amazing already.

Run Windows 2000 on a DEC Alpha with a new es40 fork

As you might know, I'm involved a bit in the [OpenVMS](/s/tags/openvms.html) community and the [Alpha emulation side via AXPBox](/s/blog/AXPBox-version-1.0.0-released.html). AXPBox ([github](https://github.com/lenticularis39/axpbox)) is a fork of the es40 alpha emulator by Camiel Vanderhoeven (who is now Chief Architect at VSI, the company that makes OpenVMS, [for x86 nowdays](/s/blog/OpenVMS_9.2_for_x86_Getting_Started.html)). There have been many forks of es40 in the past and recently [a new one](https://github.com/ES40-Emu/es40) has popped up with some great new features. Like speedups via a JIT compiler, S3 graphics port from MAME and ARC support, resulting in the ability to run Windows 2000 for the DEC Alpha.

Recently

Street art in Porto, probably commentary on tourism

June was a big month: I went to Porto & Lisbon, and had a lot of life stuff happen. I'll get into the trip once I get my rolls of film developed. Three rolls at a new photo developing place: fingers crossed!

Reading

I finished reading Intermezzo (of the bag) and it was fantastic. I've always liked Sally Rooney's books but this was the one where the writing style really clicked.

Also, The Vegan. Meh.

I read Patricia Lockwood's 'A Tradcath Wedding' via Perfect Sentences but found an additional sentence to be perfect:

Whenever they rang the chimes, which seemed to be every four seconds or so, a toddler screamed ‘WOW A BELL!’ to the visible displeasure of the celebrants – though isn’t the entire point of the ritual that you’re supposed to be that awestruck every time?

Patricia Lockwood is the funniest writer I've read.

You cannot grow a pumpkin, but you can improve the odds.

Taylor's 'You Cannot Grow a Pumpkin' is a fantastic little prose poem of sorts.

Watching

I'm always trying to find a 'romp' when it comes to movies. Something lighthearted, pretty easy to watch, so on. We watched The Pink Panther this month and it is a perfect example of the genre. The inspiration for the watch came from the hamburger scene:

But there's so much more of this kind of thing in the movie, little bits and physical comedy.

Oh, and I also watched The Departed, which everyone says is good and is good.

Elsewhere

I wrote Accidental Anonymity on the micro blog, and it stirred up some discussion on Bluesky as well as at least one blog response.

It was kind of an angry piece, as I said at the start. I will keep trying to stay out of the trap of writing about that topic all the time.

Listening

The only album I bought this month was Songs of Her's by Her's, which is fine. I wish I had something more profound to say about it given how the band met an unbelievably tragic end.

Maybe more influential than that was Know Your Enemy's recent podcasts, especially this one about the pope's encyclical. I've really grown to love that podcast, and it has been part of me intellectually reconnecting with, but not readopting, Catholicism.

Art

Here's some art I really liked this month:

David Hockney's "Picture Emphasizing Stillness"

This Hockney piece called 'Picture Emphasizing Stillness' from 1962 was at the MAC/CCB museum in Lisbon.

Tomas Sanchez

Via Tim Babb, I enjoyed finding Tomás Sánchez's work.

Two month calendar format for pocket notebooks

Hello RSS readers! Project Inbox 2026 is coming along very nicely, thank you. Don't worry, that's still my only MAIN project. I do other small things in parallel with the main project. This entry is a small thing I did recently in my pocket notebook "system" that I'm very pleased with and would like to share...

EPIPE on write might mean you're doing it wrong

Last month, I had an opportunity to dip into a part of the world I don't normally touch: Apache (as in, the web server) and the way it runs PHP code. This might seem ironic to some people since I used to support a colossal amount of PHP-ish code, but that was done with a virtual machine and had long since evolved out of the Apache ecosystem.

Yep, this is about Apache, php-fpm, mod_fastcgi, and all of that other stuff, and it was all new to me. The question was basically: why is this getting plugged up sometimes? The rest of the machine seems fine, so why is this one part going stupid?

There was a lot of random badness I ended up tripping over, but there was one part in particular I wanted to call out for the benefit of anyone who still gives a shit about trying to build this stuff properly. It has to do with being aware of the situation and not smacking yourself in the head with a shovel if you can help it.

In the above regime, a request comes in, and the web server pushes it down a Unix domain socket to a bunch of PHP worker threads which are hanging out in accept(). One of them "wins" and is rewarded with a new file descriptor for the incoming request.

The next thing it does (and this is fine) is to call poll() because it wants to wait until there's some work to be done. It immediately returns, but it has two flags set, not one. It has POLLIN, sure, but it also has POLLHUP which basically means that the far end has already gone away. It doesn't notice this second flag.

What does it do? If you said "it reads the request and runs it anyway, then sends a response to an uncaring socket that has nobody on the other end", you're right. It gets EPIPE on the write, and presumably some other part of the process gets a SIGPIPE, assuming it hasn't squashed them in its signal setup.

Let's say these requests are relatively expensive and take up to 5 seconds to run. Maybe there are 300 of them stuck in the pipe, having already been transmitted by the web server, but both it and the original http client are long gone. The fpm workers are still going to pop each one off in turn, will blow multiple seconds doing the work, and will then reply to nobody in particular.

It'll take at least 1500 worker-seconds for it to dig out from this backlog. If there are 10 workers, well, that's at least 150 seconds of it being completely saturated and thus unable to run PHP stuff for any part of the site which ends up on these workers. Good times.

This is what it looks like when this happens with multiple threads:

1672855 01:09:05.333624 write(5, "\1\6\0\1\0A\7\0Content-type: text/html; charset=UTF-8\r\n\r\nstuff... more stuff...\n\0\0\0\0\0\0\0\1\3\0\1\0\10\0\0\0\0\0\0\0\0\0\0", 96) = -1 EPIPE (Broken pipe) <0.000028>
1672853 01:09:05.335072 write(5, "\1\6\0\1\0A\7\0Content-type: text/html; charset=UTF-8\r\n\r\nstuff... more stuff...\n\0\0\0\0\0\0\0\1\3\0\1\0\10\0\0\0\0\0\0\0\0\0\0", 96) = -1 EPIPE (Broken pipe) <0.000027>
1672854 01:09:06.307556 write(5, "\1\6\0\1\0A\7\0Content-type: text/html; charset=UTF-8\r\n\r\nstuff... more stuff...\n\0\0\0\0\0\0\0\1\3\0\1\0\10\0\0\0\0\0\0\0\0\0\0", 96) = -1 EPIPE (Broken pipe) <0.000028>

That's my dumb little test script which prints "stuff... more stuff..." being run by the fpm worker and finding out that the requester is gone. This was many minutes after I stopped sending new requests.

This is how you get a system that will stay sick long after the crushing load of requests has gone away. Again, it could have known this at the very beginning...

01:20:36.514261 poll([{fd=5, events=POLLIN}], 1, 5000) = 1 ([{fd=5, revents=POLLIN|POLLHUP}]) <0.000016>

It's right there! POLLHUP! It's gone! The poll loop in question is itself buried maybe six or seven levels deep in a massively nested set of while loops, if-this, if-that, etc, and amounts to this:

do {
  errno = 0;
  ret = poll(&fds, 1, 5000);
} while (ret < 0 && errno == EINTR);

if (ret > 0 && (fds.revents & POLLIN)) {
  break;
}

It just sits there in poll for 5 seconds, ignoring the usual worse-is-better EINTR stuff, until poll says something is ready to rock or it gets a "real" error, or it times out.

But if poll said something went active and POLLIN is set, then it breaks out of whatever "while (1)" loop it's in (seriously), and carries on to where it reads from the socket and fires up the parser and runs some script with the parameters. That's what you'd expect.

I wondered what it would do if I forced it to abort any request where POLLHUP was set on the client fd. It should just drop straight through and clean up the mess relatively quickly, and sure enough, it did:

3756041 04:09:59.049764 poll([{fd=5, events=POLLIN}], 1, 5000) = 1 ([{fd=5, revents=POLLIN|POLLHUP}]) <0.000016>
3756041 04:09:59.049827 close(5)        = 0 <0.000024>
3756041 04:09:59.049888 accept(10, {sa_family=AF_UNIX}, [112 => 2]) = 5 <0.000030>
3756041 04:09:59.049972 fcntl(5, F_GETFD) = 0 <0.000013>
3756041 04:09:59.050021 fcntl(5, F_SETFD, FD_CLOEXEC) = 0 <0.000013>
3756041 04:09:59.050070 poll([{fd=5, events=POLLIN}], 1, 5000) = 1 ([{fd=5, revents=POLLIN|POLLHUP}]) <0.000015>
3756041 04:09:59.050130 close(5)        = 0 <0.000023>
3756041 04:09:59.050190 accept(10, {sa_family=AF_UNIX}, [112 => 2]) = 5 <0.000019>
3756041 04:09:59.050251 fcntl(5, F_GETFD) = 0 <0.000013>
3756041 04:09:59.050300 fcntl(5, F_SETFD, FD_CLOEXEC) = 0 <0.000013>
3756041 04:09:59.050349 poll([{fd=5, events=POLLIN}], 1, 5000) = 1 ([{fd=5, revents=POLLIN|POLLHUP}]) <0.000014>
3756041 04:09:59.050407 close(5)        = 0 <0.000021>
3756041 04:09:59.050465 accept(10, {sa_family=AF_UNIX}, [112 => 2]) = 5 <0.000019>
3756041 04:09:59.050523 fcntl(5, F_GETFD) = 0 <0.000024>
3756041 04:09:59.050579 fcntl(5, F_SETFD, FD_CLOEXEC) = 0 <0.000007>
3756041 04:09:59.050616 poll([{fd=5, events=POLLIN}], 1, 5000) = 1 ([{fd=5, revents=POLLIN|POLLHUP}]) <0.000011>

Just... bonk, bonk, bonk, bonk. The first worker that wakes up ends up stripping all of these useless requests from the queue. It doesn't spend any time reading them, never mind parsing them, and it definitely doesn't try running any PHP code. It just closes the fd and goes back to waiting for another connection.

It managed to destroy a backlog of a few thousand dead requests in under a second. By not wasting time on those things, it became available to handle the real requests from the clients who were still connected to the web server right then.

Now, is this a universal solution? Probably not. There are plenty of async situations where you'll get POLLHUP and you still need to read until EOF because you care about consuming whatever data managed to make it down the chute before the far end hung up. Case in point: do the fork, dup2, exec thing to run a subprocess and watch the fds for stdout and stderr from the child process. poll will totally see POLLHUP once it's done, but you have to stick around to consume the rest of it. You did want the actual output from the process, right?

But, if you're talking about individual webshit requests that are probably going to be retried anyway and which don't need to be executed "at any cost" because there's nothing unique or special about them, you could probably stand to slough off the load up front.

Bypassing the Sound Blaster's new firmware signature check

I recently demonstrated a few vulnerabilities in the Sound Blaster Katana V2/V2X/SE which allowed me to hijack the device over Bluetooth and turn it into an attacker-controlled keyboard peripheral, injecting keystrokes into the connected machine.

Creative recently published a new firmware (version 1.20.260617.1110) to address these issues. In this article, I take a deep dive into the changes and highlight why they're not sufficient to stop the same attack that I described.

Factories are just rooms

I went into my kid’s school a couple months back and spoke to the year group about manufacturing.

Honestly it was the most rewarding speaking gig I’ve done all year.

It was about the process of making my AI clock and I have a ton of pics from my factory visit to Shenzhen (mostly pics that I have only shared with Kickstarter backers).

I talked about where ideas come from and the value of playing around, and how it’s neat to learn new techniques that you can combine together.

I talked about prototyping and design – and was sure to use the words “prototyping” and “design”. I showed exploratory sketches and what CAD looks like.

I handed round various iterations of e-paper screens, and electronics from breadboard to PCB, and various iterations of plastic parts.

It’s interesting to see how a plastic enclosure comes apart, and to connect that to what an injection moulding machine is doing.

(A lot of the kids are familiar with 3D printers, so I showed a timelapse of a 3D print – it would take a year to print all my clocks! And then a real-time video of injection moulding, and how that would only take a day.)

And then photos of factory floors, and here’s the team, and assembly lines and what a page from an assembly procedure looks like, and packaging too.


7-year-olds have great questions.

Like: how does it not break in the post?

Well here’s a vibration machine in action and that’s how we test it.

And, look, in this cardboard packaging, here’s a cradle, and this was made by a packaging designer – you could be a packaging designer too if you want.

Like: how does the button work?

Well you’re right I didn’t pass round the separate button piece, good spot, it’s small and I didn’t want to lose it. So now let’s talk about assembly and about industrial designers…


I don’t like those videos of factories that are supposed to inspire awe.

You know the ones I mean: you see a thousand products a second whizzing by on 20 parallel belts. You come away saying wow. When they showed manufacturing on kids’ TV when I was growing up, that was what they showed.

“Awe” is the opposite of what I want to convey.

Except for a very specific type of person, when you show something with the expectation that “awe” is the appropriate response, you are implicitly saying to your audience: you should step back here and appreciate this from a distance. Like looking at a great work of art. Gasp but do not place yourself in the picture.

Whereas!

I want to re-home manufacturing. I want these kids to become designers, engineers, inventors, factory owners, and all the rest. Makers of any kind; participants in the ongoing making of our world.

So my message is: sure this is complicated but it’s fine, we can do complicated.

Factories are just rooms.

The stuff around us isn’t divine - these chairs we’re sitting on, the TV at the front of the classroom, the pots for the plants - all this stuff was invented and figured out and made by people.

p.s. you can be one of those people.


So when I heard the class was learning about inventing, I offered to go in and show that scrappy dead ends are cool actually (it was amazing to speak with a class that already knows the word “prototyping”) and this is electronics and this is going from sketching to plastic and this is what it means to make a product and to sell it.

I deeply feel this mission to normalise getting our hands dirty with the world – when they’re 7 years old, while their brains are still establishing what’s normal.

(This is connected with what I was saying about training for collective efficacy.)

And I’m just someone’s dad, you know? So if this guy can do it…

If you have the opportunity to go into your local school and talk about making things too, please do. You will be rewarded with wonderful curiosity, engagement and questions from the kids.

Hopefully one of them one day will look around them, think “someone should do something about that”, remember back, and say - oh that someone can be ME.


More posts tagged: that-ai-clock-and-so-on (15).

Auto-detected kinda similar posts:

Filtered for that which motivates form

1.

It’s hot in London so I’ve been seeing a lot of handheld portable fans, usually with a strap so you can hang it round your neck.

My faves are the ones with thermoelectric coolers in the middle of the fan: a small plate that is so cold that it is covered with icy condensation. It’s the Peltier effect (the plate is really hot on the back) and the first time I’ve seen a thermoelectric cooler in the wild.

Here’s a thermoelectric handheld fan on Alibaba. Three quid each if you’re buying over a million.

A fan looks like a fan because of the mechanism used for the movement of air.

I’ve been meditating this week on what motivates form in product design.

Hey free concept: AirPods with built-in Peltier thermoelectric coolers so the buds are ice-cold in your ear holes.

2.

A fan moves air; electronic products move data. That can motivate the form just the same.

Durrell Bishop’s Marble Answer Machine (watch the video):

marbles dropping out of an answering machine could form an intuitive physical interface. This work later became the seed for a new movement called Tangible User Interfaces.

Each marble is a message that you can place in the player dish, put aside to keep for later, and so on. So sophisticated and so immediately understandable. (“Legible” as Durrell say.)

But what about when a product can do anything? Like a phone?

We end up with anonymous slabs of black glass.

3.

Before movie theatres there was the Kinetoscope (Wikipedia) "an early motion picture exhibition device, designed for films to be viewed by one person at a time through a peephole viewer window."

The Kinetoscope came out of Thomas Edison’s lab and established the idea of reels of film, and also film as “content” to be manufactured and distributed.

But it was a single-viewer device: you leant down and put your eyes to the viewer.

The act of peeping – the form is motivated by the human interaction.

Or for the necessity of the affordance: a possible interaction that must been seen as a possible interaction. (Affordances as previously discussed.)

BTW:

I recently discovered via the sf journal [Foundation] (issue 152) that:

Thomas Edison, using his Kinetoscope, is credited with producing (though not directing) … the first filmed sneeze (1894), the first filmed kiss (1896).

4.

The egg timer that looks like an egg is the best product design of all time.

Here’s one on Amazon.

The traditional kind of hourglass timer that uses sand, on the other hand, is rubbish:

  • It is traditional
  • Its form is motivated by the mechanism
  • And it is legible because you can see the passing of time.

But what to do you use it for?

Measuring time, sure. A lot of stuff. It can do anything (related to waiting for a period of time).

But does an initial use come to mind?

You have to think about it – aha eggs! Or you have to learn it. Ultimate the sand timer is abstract.

Like an empty ChatGPT window?

AI can do anything too.

But I opened my ChatGPT just now, and look how hard they work to give you ideas of what to do. Mine suggests: Write an email, create a painting, give me ideas…

If the hourglass timer were designed like that, they would print on the side:

  • For eggs
  • For rice
  • For pasta
  • For reminding me how long I have till my program is on.

So cumbersome.

Whereas!

The egg timer that is shaped like an egg.

Form follows function – but also where and when and how.

And then once you have timed your eggs, you have in your mind this new hammer of “timing” and you see immediately everything else you can time.

So another motivation for form is to imply the first context of use, even if - and especially if - the product can be re-purposed or adapted for other contexts by the end user, once they have internalised the function of it.

It is genius. I aspire to design a product this perfect.

The egg timer shaped like an egg was invented and patented by Lucio Oliveri in 1982. U.S. Design Patent No. D276,705 (expired).


More posts tagged: filtered-for (124).

Examining circuit boards from the Space Shuttle's I/O Processor

The Space Shuttle's five1 general-purpose computers played a critical role in each flight: controlling the engines, monitoring thousands of sensors, displaying data to the astronauts, and navigating the Shuttle. Each computer consisted of two 60-pound aluminum-alloy boxes: the box on the right is the CPU, a 32-bit processor that executed 420,000 instructions per second. These computers were designed before microprocessors became popular, so the processor was built from multiple boards crammed with simple chips and they used magnetic core memory rather than DRAM chips.

The Space Shuttle IOP and CPU (AP-101B). Photo courtesy of RR Auction.

The Space Shuttle IOP and CPU (AP-101B). Photo courtesy of RR Auction.

The box on the left is the I/O Processor (IOP): the link between the CPU and the rest of the Shuttle. It implemented the input/output capabilities for the computer, primarily 24 high-speed networks that connected the computer to the Shuttle's systems and sensors. But the IOP wasn't just a peripheral; it was a separate programmable computer, more complicated than the main CPU. The IOP had an unusual architecture: it was one of the first multi-threaded computers, implementing 25 virtual processors (with two completely different instruction sets) that ran on one physical processor.

I obtained two circuit cards from the I/O Processor,2 each a 9"×3" rectangle packed with tiny chips and other components. In IBM lingo, each card is called a "page" (remember this term). The top page is a network interface, providing four network connections, each handling 1 million bits per second. (The IOP contained six of these cards for its 24 network connections.) The bottom page held the microcode for the IOP's processors, the low-level code that defined each instruction. The rows of white-and-gold chips stored the microcode's bits in tiny metal fuses, programmed by blowing a fuse for each 1 bit. In this article, I'll explain how the I/O Processor worked, and the roles of these two pages.

Two pages from the Space Shuttle I/O Processor: the "MIA" interface page and the PROM page.

Two pages from the Space Shuttle I/O Processor: the "MIA" interface page and the PROM page.

The MIA interface page

The Space Shuttle had 28 data bus networks that linked the computers to the rest of the Shuttle, with each computer attached to 24 of the networks.3 The large number of networks provided both high performance and reliability, with at least two networks between a computer and any Shuttle system. Eight networks were assigned to flight-critical systems, with each CRT display and engine controller connected to four networks for redundancy.

The page below is one of the six network interface pages in the I/O Processor. Space Shuttle engineers loved acronyms, so this page has the cryptic name MIA for "Multiplexer Interface Adapter". (Many of the networks were connected to boxes called Multiplexer/Demultiplexers, which provided the link between the network and the diverse analog and digital components of the Space Shuttle.5) The MIA interface page is tightly packed with integrated circuits and other components. The page holds two printed-circuit boards, one on each side of the page. The boards on both sides are almost identical,4 as you can see by comparing the photo above and the photo below. (Main difference: the connector switches sides.)

The network interface page, called the MIA (Multiplex Interface Adapter).
The page has extensive rework; thin brown "bodge" wires snake around the page to
repair errors or implement updates.

The network interface page, called the MIA (Multiplex Interface Adapter). The page has extensive rework; thin brown "bodge" wires snake around the page to repair errors or implement updates.

Each board implements two network interfaces, so the page supports four networks. Each network transmits data across a pair of wires, twisted together and shielded, rather than a coaxial cable. Although the network transmits digital data, the signals transmitted across the network are physical voltages that will weaken with distance and will have distortion and noise. Thus, the interface page must convert these analog signals back to 0's and 1's.

The right half of the board holds the analog circuitry. It is dominated by a large golden module labeled "IBM", with 46 pins. This is a hybrid module, consisting of tiny components such as transistor dies, resistors, capacitors, and potentially IC dies, connected by bond wires thinner than a hair. It's not quite an integrated circuit, but a collection of individual components mounted on a ceramic wafer. Hybrid modules were popular for aerospace applications, since a board of analog components could be shrunk down to a single (expensive) module. This module contains the analog circuitry for two I/O ports: the drivers to transmit network signals along with the amplifiers and comparators to receive signals.

Various discrete components are mounted next to the hybrid module: resistors, glass capacitors6, inductors, and small square transformers. The transformers provide the coupling between the interface board and the network. As with Ethernet, transformers provide isolation between the computer and the network, filter electromagnetic interference, and match impedances, all important for reliability.7

The Manchester Mark 1; Prof. Williams is second from the left. Photo from the University of Manchester.

The Manchester Mark 1; Prof. Williams is second from the left. Photo from the University of Manchester.

A key part of the Shuttle's networking dates back to the 1940s. In 1946, Frederic Williams became head of the Electrical Engineering department at the University of Manchester. By 1949, his team had created the groundbreaking Manchester Mark 1 computer. Along the way, they invented the stored-program computer, the Williams tube—the best form of computer memory before magnetic core—and the Manchester Carry Chain, still used for addition in modern processors.

But the relevant invention is the patented Manchester encoding, a way of encoding a sequence of 0's and 1's for storage or transmission. In the Manchester encoding, each 0 bit is replaced by a "low-high" sequence and each 1 bit is replaced by a "high-low" sequence, as shown below. This idea may seem trivial, but it is used in everything from floppy disks and remote controls to Ethernet and RFID tags, earning it recognition as an IEEE Milestone.

A diagram illustrating Manchester encoding. From Prototype IOP Functional Description, p82.

A diagram illustrating Manchester encoding. From Prototype IOP Functional Description, p82.

The obvious approach—sending binary data unencoded—has two problems. First, in a long string of 0's or 1's, it is hard to tell how many bits were sent: "Was that six bits or only five?" Second, such a sequence is unbalanced, so it has a "DC component". This DC component causes problems if the signal is stored on a magnetic medium or transmitted through a transformer. The Manchester encoding solves both these problems. Since every encoded bit has a transition in the middle, it is straightforward to separate the bits. Moreover, the encoding ensures that 0's and 1's occur in equal numbers, so there is no DC component.

Because of these advantages, the Manchester encoding was selected for the data bus networks in the Space Shuttle.8 One of the key functions9 of the IOP's network interfaces is to convert between serial bits and the Manchester encoding. The digital circuitry for the interface is fairly complicated, but most of the logic is in the four large golden integrated circuits. These are custom Motorola integrated circuits: a transmit chip and a receive chip for each network port. On the transmit side, the chip converts binary data into the Manchester-encoded signals for the network. The circuitry also inserts a sync signal at the beginning of each word and adds parity. The receive chip reverses this process: detecting sync, decoding the Manchester signals, verifying the parity, and reporting any errors.

The smaller black chips are simple TTL chips, mostly shift registers. (Transistor-Transistor Logic was very popular in the 1970s, providing fast, reliable circuits.) There are twelve 4-bit shift register chips and sixteen 8-bit shift registers.10 The Shuttle's networks sent 24-bit words across the network: combining six 4-bit shift register chips produces a 24-bit shift register, which converted these 24-bit words to serial data and vice versa. The remaining chips are simple logic gates, flip-flops, buffers, and four-bit counters.

The physical structure of a page

Around 1967, IBM introduced a line of computers for avionics, called System/4 Pi.11 These systems were constructed from pages:12 two circuit boards sandwiching a metal layer that provided conduction cooling. Flat-pack integrated circuits, smaller than a fingernail, were mounted in rows13 on each circuit board, about 78 ICs on a board. The printed-circuit boards were advanced for the time, with six layers of wiring. Two jack screws at the top tightly secured the page into the system. Two 98-pin connectors connected the page to the backplane. The photo below shows a typical 4 Pi page (top), with its rows of chips.

A comparison of a standard IBM 4 Pi page with the IOP page. 4 Pi page courtesy of Eric Schlaepfer. The 4 Pi page was in a bag labeled "FSD AWACS tester?" suggesting that it was a tester from IBM's Federal Systems Division for the E-3C Airborne Warning and Control System aircraft, which used an IBM 4 Pi computer.

A comparison of a standard IBM 4 Pi page with the IOP page. 4 Pi page courtesy of Eric Schlaepfer. The 4 Pi page was in a bag labeled "FSD AWACS tester?" suggesting that it was a tester from IBM's Federal Systems Division for the E-3C Airborne Warning and Control System aircraft, which used an IBM 4 Pi computer.

An I/O processor page (above, bottom) is almost identical to a standard 4 Pi page except that it is one inch wider (9" instead of 8"), and has a 120-pin connector or two instead of 98-pin connectors.14 One inch may not seem like much, but a 9-inch page fits 100 ICs rather than 78, a significant increase. I'm surprised that IBM changed from the standard size, but I suspect that the designers couldn't fit the IOP into the available space with standard pages, forcing the change. Likewise, the multiple I/O ports may have required more connections than the smaller connectors could support.

A page has circuit boards on either side, separated by a metal plate. To allow signals to flow between the boards, a special connector is attached to the top of the page to link the two boards. This connector not only provides feed-through connections between the boards, but also provides test points, so signals can be probed while the boards are mounted in the case. The photo below shows a close-up of the feed-through connector. It has three rows of test points. The first row (red) is connected to the top board. The middle row (orange) is connected to both boards and provides the feed-throughs. The bottom row (blue) is connected to the bottom board. The upper arrows show where the connector is soldered to the board.

The test point connector on the MIA page.

The test point connector on the MIA page.

The diagram below shows the construction of the I/O Processor, with rows of pages plugged into the backplane.15 Note the 128-pin MIA I/O connector on the front of the IOP; this connects the 24 data buses (along with other signals) to other parts of the Shuttle. The arrows show how cooling air flowed through the sides of the IOP. The air did not flow over the pages. Instead, heat was transmitted by conduction through the metal plate inside each page, flowing to heat exchangers in the sides of the case. The CPU and the IOP both contained magnetic core memory (labeled "Storage Page" below); even though the memory is split between the boxes, it is treated as a unified shared memory, so programs for the CPU and the IOP can reside in memory in either physical box.

Exploded view of the IOP. From Prototype IOP Functional Description.

Exploded view of the IOP. From Prototype IOP Functional Description.

The IOP's architecture and the PROM page

The high-performance design of the I/O Processor was developed by Peter Kogge, an expert in parallel processing architectures. At the time, he was working at IBM's Federal Systems Division, where the Space Shuttle computer was developed.24 Kogge, now a professor at the University of Notre Dame, is also known for the Kogge-Stone adder, a fast circuit used in processors such as the Pentium. The I/O Processor has a very unusual architecture: although it had one physical processor, it ran 25 virtual processors with two completely different instruction sets. The virtual processors took turns, running for just one clock cycle and then letting the next processor run. The motivation behind this was to ensure that each network port got a predictable and guaranteed portion of the processor, so even if one network port was overloaded, it wouldn't affect the others. This approach, called a barrel processor16, was first used in the CDC 6600 supercomputer, the world's fastest computer from 1964 to 1969.

The I/O Processor has two types of (virtual) processors, which of course have cryptic acronyms: BCE and MSC. Each of the 24 network ports has a BCE, a Bus Control Element, which runs a small program to move data words between the network port and memory. An MSC (Master Sequence Controller) is the executive, running programs to manage the BCEs. The BCE and MSC processors run code that is stored in the computer's core memory. The instruction sets of the MSC and the BCE are completely different from each other and from the instruction set of the main CPU (which is derived from IBM's System/360 mainframes). The (executive) MSC is a 32-bit processor with the standard instructions of a normal processor—addition, logic, branches, and so forth—as well as specialized operations to configure and start BCEs.17 The instruction set of a low-level BCE is much smaller and much stranger, lacking all the basic instructions such as arithmetic and conditional branches. the instructions you'd expect from a processor. Instead, a BCE has I/O instructions such as Transmit Data, Receive Data, Load Timeout Register, Store Status, and Wait. In typical use, the CPU directs the MSC to run a program, the MSC configures the BCEs to execute a program, and the BCEs send and receive data as specified. When the BSE's operation is done, the MSC interrupts the CPU, which processes the data. Thus, the CPU can focus on the high-level algorithms without wasting cycles on network operations.

How do the MSC and BCE processors all run on one physical processor, when they have completely different instruction sets? The trick is microcode: each MSC and BCE instruction was implemented in microcode, through a sequence of 72-bit micro-instructions.18 A simple instruction might take five micro-instructions, while a complex instruction might require 60 micro-instructions. Each micro-instruction directed the action of the IOP's physical processor for one step of the MSC or BCE instruction. After each micro-instruction, the physical processor switched to the micro-instruction for the next virtual processor. The architecture of the physical processor was completely different from the MSC or the BCE: three 16-bit data paths and two ALUs (Arithmetic/Logic Units) that can operate in parallel. The physical processor had a separate register set, including a micro-instruction address register, for each virtual processor, to keep track of the state of each virtual processor.

The PROM page holds the majority of the microcode for the I/O Processor. Although three chips are mounted sideways to avoid wasting space, there is even more wasted space at the left.

The PROM page holds the majority of the microcode for the I/O Processor. Although three chips are mounted sideways to avoid wasting space, there is even more wasted space at the left.

The IOP's micro-instructions were stored in the PROM page above. In the photo above, the white chips with gold lids are fusible-link PROM (Programmable Read-Only Memory) chips.19 These unusual chips contain a tiny fuse for each bit. If the fuse is intact, the corresponding bit is a 0, while a burnt-out fuse represents a 1 bit. The chip is programmed by applying 17-volt pulses to destroy fuses one by one, literally burning the PROM. (I discussed fusible PROM chips earlier.)

Each PROM chip holds 512 words of 4 bits, so in total, this page held 1024 72-bit micro-instructions; the remaining 512 micro-instructions were in another page.20 The chips are hand-labeled with numbers, since each chip has unique programming and must be installed in the correct location. With 36 chips, you'd expect the chips to be numbered from 1 to 36. Curiously, although many of the chips are sequentially numbered, others have numbers ranging from 55 to 74 in no obvious pattern.21

Physically, the PROM page is unusual in several ways. Instead of flat-pack integrated circuits, it uses DIP (Dual-Inline Package) ICs, larger integrated circuits with two rows of vertical pins that go through the circuit board. Since this page only has one circuit board, it doesn't have the test-point feed-throughs at the top. It still has the central metal plate, but the integrated circuits sit on top of the metal plate, while the circuit board is underneath—the plate has gaps for the pins. Between the rows of chips, the central plate is the full thickness of the board.

A close-up of the PROM page, showing how the chips are mounted. The black chips are much thicker than the white chips.

A close-up of the PROM page, showing how the chips are mounted. The black chips are much thicker than the white chips.

Presumably, the fusible-link PROM chips were only available in DIP packages, rather than flat-packs. These DIP packages take up much more space than the regular flat-pack integrated circuits; this page has about a quarter the density of a regular page.22

Conclusions

The Space Shuttle's CPU and IOP were advanced when they were designed, but they rapidly became obsolete. IBM redesigned the computer, combining both the CPU and IOP into a single box called the AP-101S, which first flew in 1991 (details). The improved computer was much faster and had more memory. Moreover, combining two boxes into one saved about 300 pounds in total. The photo below shows three of the updated AP-101S computers mounted in the Shuttle's avionics bays. (The wall hides the fourth computer, and the fifth is behind the camera.) These same positions are where the I/O Processors were mounted previously, with the CPUs installed in the empty spaces to the left.

Avionics bays 1 and 2 are located in the crew cabin middeck, below the flight deck, and looking forward into the nose. The red arrows indicate the AP-101S computers. The remaining computer is in avionics bay 3A, on the aft right side of the middeck. This photo is from 2011, showing Discovery being prepared for display at the Smithsonian. Original photo courtesy of collectSpace; I've adjusted the lighting.

Avionics bays 1 and 2 are located in the crew cabin middeck, below the flight deck, and looking forward into the nose. The red arrows indicate the AP-101S computers. The remaining computer is in avionics bay 3A, on the aft right side of the middeck. This photo is from 2011, showing Discovery being prepared for display at the Smithsonian. Original photo courtesy of collectSpace; I've adjusted the lighting.

Despite the critical role of the I/O Processor in the Space Shuttle, it doesn't get the attention given to the CPU. For instance, although NASA documents describe the architecture of the IOP in detail, I couldn't find any photos of its pages.23 I hope that this article has convinced you that the architecture and the physical construction of the IOP make it an interesting system.

For updates, follow me on Bluesky (@righto.com), Mastodon (@kenshirriff@oldbytes.space), or RSS. Thanks to Richard for supplying the boards. Thanks to Mike Stewart for documents on the IOP. Thanks to Robert Pearlman of collectSPACE, and RR Auction for photos.

AI statement: I didn't use AI to write this article; the em-dashes are natural (details).

Notes and references

  1. On some flights, a sixth computer was carried in a locker as a spare, providing an additional degree of reliability. If one of the five computers failed, the astronauts could connect the cables to the spare computer and it could take over for the failed one. The spare was put into use on flight STS-30 (1989) after computer #4 encountered a "data parity external storage error", indicating a hardware problem. 

  2. I suspected that these pages were from the I/O Processor, but it was difficult to prove this. Fortunately, Mike Stewart found a document, the Prototype Input/Output Processor Function Description, that lists the pages in each IOP slot. The MIA page has a part number on it: 6246523-3, and the PROM page has 6104848-3; these match "MIA" 6246523-1 and "Micro Store (ROM)" 6104848-1 in the document. 

  3. The diagram below shows how the 28 data bus networks connect the five computers at the top and various parts of the Shuttle. The networks are categorized as ground interface, mission critical, flight instrumentation, display system, mass memory, intercomputer, and flight critical.

    Data bus architecture. Click for a larger version. Adapted from Space Shuttle Avionics Systems.

    Data bus architecture. Click for a larger version. Adapted from Space Shuttle Avionics Systems.

    Why was each computer connected to 24 networks and not all 28? Each Space Shuttle computer was connected to almost all the networks, so they could run in lockstep for reliability. The exception was that each computer sent its own monitoring data to the ground station. Since this data was of no importance to the other computers, it was sent over a private network called Flight Instrumentation to the PCM (Pulse Code Modulation) box, which encoded the data for transmission to the ground. There were 23 shared networks and 5 private networks (one for each computer), so there were 28 networks in total, with 24 networks connected to a particular computer. 

  4. Both sides of the interface page are almost identical. However, the connector is on the left or the right side, depending on which side of the page you examine. This forced the decoupling capacitors at the very bottom to move to accommodate the connector. I also found a single integrated circuit that was different between the two sides, for some reason. 

  5. While many of the data bus networks are connected to a Multiplexer/Demultiplexer (MDM), this is not always the case. Networks were also connected directly to systems such as an Engine Interface Unit or a Display Electronics Unit. Moreover, the MDM was not necessarily the final step between the network and the Shuttle's sensors. The MDM held cards to support over a dozen types of input and output signals: digital, analog, on/off (discrete), and serial. However, the thousands of signals in the Shuttle were much more diverse; sensors can provide AC signals, pulses, thermocouple values, resistances, and so forth. Other boxes converted the raw sensor signals into forms that the MDM could handle; these boxes were called Dedicated Signal Conditioners (DSC). A DSC had 15 or 30 slots to hold cards to perform the necessary signal conversion. Thus, the MDMs and DSCs combined a fixed architecture with the ability to be customized for each role. 

  6. The glass capacitor is an interesting component, with an extremely thin layer of glass as the dielectric. Glass capacitors became popular in the 1960s for aerospace applications because of their stability and reliability (more). These capacitors were manufactured by Corning Glass Works, as indicated by the "CGW" label on the package.

    Two glass capacitors on the MIA page.

    Two glass capacitors on the MIA page.

    The capacitor is labeled with a military code. "J" indicates the Joint Army/Navy specification. "CY" indicates a glass capacitor, "4" apparently indicates axial leads, "G" indicates the temperature/voltage, "510" is the value (51×100 = 51 pF), and "G" indicates ±2% tolerance. (I don't know why one capacitor has "0F" and the other has "4G".) 

  7. The Space Shuttle had a second layer of transformers between the computer and the network, ensuring a faulty device didn't bring down the network. Each device (such as the IOP) was connected to the network through a tiny device called the Data Bus Coupler. This one-inch cube contains a transformer and a few resistors to match impedance. The coupler acts as a network tap, providing a short stub from the network to a device. The coupler also provides line termination if the device is removed, ensuring signal integrity. 

  8. The Space Shuttle's network is very similar to the U.S. military's serial network standard MIL-STD-1553. The 1553B standard is widely used in numerous military aircraft, missiles, tanks, navy systems, the Airbus A350 commercial plane, and the James Webb Space Telescope. However, since the Space Shuttle's network and the 1553 standard were both under development in the early 1970s, the two networks are not the same. The main differences are that the Shuttle uses 24-bit words instead of 16, and has 5.5µs gap between words (details). 

  9. The functions of the MIA are described as:

    • Transmit and receive data
    • DC isolation
    • Parallel/serial conversion
    • Serial/parallel conversion
    • Sync generation and detection
    • Manchester encode and decode
    • Parity generation and detection
    • Bit count detection
    • Provide status to BCE.

    The functional block diagram below shows the circuitry for one port of the network interface. This circuitry is replicated twice on each board; with a board on each side of the page, the page supports four networks. The dashed Transmitting and Receiving boxes correspond, I think, to the large Motorola chips, except that the "TX" and "RX" amplifiers are in the IBM hybrid module and the transformers are discrete components.

    Functional block diagram of the MIA. From Prototype IOP Functional Description, p82. Click for a larger image.

    Functional block diagram of the MIA. From Prototype IOP Functional Description, p82. Click for a larger image.

     

  10. The 4-bit shift register chips are 54LS395 chips. These chips have "tri-state" outputs, allowing them to be connected to a bus. These chips probably provide the interface between the board and the rest of the IOP; the twelve chips on a board would support a 24-bit register for each port, as expected. The 8-bit shift register chips are 54LS1964 shift registers.

    I can't figure out why there are so many 8-bit shift register chips; perhaps they act as buffers. My speculation... The Prototype IOP Functional Description states that the IOP has six 28-bit 4-word registers between the 24-bit MIA shift registers and the rest of the IOP. Could the 8-bit shift register chips form these registers, even though shifting is not necessary? The document doesn't make it clear if these registers are on the MIA page or a different page. The shift-register chips provide 256 bits of storage per page, while the register file needs 112 bits, so there are way more bits than required. Moreover, the document says that the registers are structured as 7-4&4 register files for each set of four MIAs, which sounds more like 54LS170 register file chips (for instance) than shift-register chips. Possibly, the design was modified from the Prototype Functional Description, and the 8-bit shift registers provide additional buffering. 

  11. The 4 Pi name is a geometry joke based on IBM's wildly popular series of mainframes, the System/360. System/360 revolutionized the computer industry with the concept of one family of computers for all applications: business and scientific. The name symbolized that System/360 covered the full 360º of applications. The 4 Pi name extended the idea of a circle to the 3-dimensional world: 4π is the number of steradians making up a full sphere. As IBM put it, "System/4 Pi also fills a sphere—the full spectrum of military computer needs—for airborne, space, or shipboard use." 

  12. The earliest 4 Pi systems (the TC line) used a different style of page, but the following computers used the standard 4 Pi pages, including the Space Shuttle's AP-101B computer. However, IBM moved to much larger pages, starting with the next computer, the AP-101C in the B-1 bomber. The Space Shuttle's upgraded computer, the AP-101S, used these larger pages. For details, see my article on 4 Pi computer history

  13. The photo below shows how the flat-pack integrated circuits are mounted on the circuit board. 16 pads are allocated to each integrated circuit; 14-pin integrated circuits "waste" two pads, while larger integrated circuits break the regular pattern. Each pad is connected to a via, a plated hole through the circuit board. These vias provide connections to wiring traces on a different layer of the circuit board; some of these traces are visible in the photo. Vias also hold the leads of through-hole components. The circuit cards in IBM System/360 mainframes used a very similar style of printed-circuit board, with a regular grid of vias. This style of board is very different from the circuit boards used in most other systems, which only had holes where necessary and routed traces less regularly. IBM's style presumably made hole drilling more efficient and was easier for automatic routing, but required thin, precise traces and multi-layer circuit boards, which were not common at the time.

    IBM's technology was highly advanced compared to consumer electronics. IBM was using six-layer printed-circuit boards and surface-mount components in the 1960s, but Apple, for instance, didn't switch to surface-mount components until two decades later. Specifically, the Apple IIGS (1986) extensively used surface-mount components, but the Macintosh SE (1987) still used entirely through-hole components a year later.

    A close-up of the IOP's PROM board.

    A close-up of the IOP's PROM board.

    The photo also illustrates how some integrated circuits are labeled with Specification Control Drawing (SCD) numbers (6088731-1) while others are labeled with standard part numbers (SN54LS151). This SCD number corresponds to a standard 54S10 NAND gate. The chips both have 1974 date codes (74xx), not to be confused with 7400-series part numbers.

    The photo below shows three different types of flat-pack ICs. The first type is most common, with leads extending from the top and bottom sides, similar to a modern surface-mount integrated circuit. The second package has a golden case. It is much smaller and thinner, with leads extending from all four sides. The third package also has leads from four sides, but is somewhat larger.

    Three types of surface-mount packages.

    Three types of surface-mount packages.

     

  14. The change in page size for the IOP is documented in Prototype IOC Functional Description, which says: "Standard 4 Pi Page Extended by Width Change from 8 to 9 inches, New Standard 120 Pin Connector".

    The photo below compares the 98-pin connector on a standard IBM 4 Pi page (top) with the 120-pin connector on the IOP page (bottom). The 120-pin has a narrower pin spacing (0.05") than the 98-pin connector (0.06"), allowing more pins in the same width. However, the 120-pin connector has more spacing between the rows of pins (0.150" vs. 0.100").

    The connectors on a standard IBM 4 Pi page (top) and the IOP page (bottom). The 4 Pi page is courtesy of Eric Schlaepfer. The slight waviness is just due to bent pins.

    The connectors on a standard IBM 4 Pi page (top) and the IOP page (bottom). The 4 Pi page is courtesy of Eric Schlaepfer. The slight waviness is just due to bent pins.

    Also note that both connectors have a peg on one side and a hollow cylinder on the other. These are used for keying, to make sure that a page cannot be plugged into the wrong slot. Each page type has a different combination; with a double connector, there are 16 possible combinations. 

  15. The exploded view shows seven MIA (interface) pages. This doesn't make sense since there are six MIA pages for the 24 network connections, as the same document lists (in Table 4-1). That table also shows one more page in total than on the exploded view. My guess is that the system was still being changed when the document was written (some entries in the table are marked TBD), resulting in inconsistencies. 

  16. The virtual MSC and BCE processors take turns executing on the IOP's physical processor. A 16.5 µs time interval is split into 33 slices: each BCE gets one time slice, the MSC gets 8 time slices, and one slice is used for BCE self-tests. Thus, the MSC gets much more execution time than a low-level BCE.

    The I/O Processor's slot timer or "wheel". Adapted from Space Shuttle Systems Handbook, 8.3.

    The I/O Processor's slot timer or "wheel". Adapted from Space Shuttle Systems Handbook, 8.3.

    Each BCE and the MSC has its own register set (called local store), so the right registers are available for each slot. The physical processor is pipelined, so there are actually four slots active at any time. 

  17. For details on the instruction sets of the MSC and BSE processors, see Prototype IOP Functional Description, chapter 2. 

  18. The IOP used a micro-instruction that was 72 bits wide. A micro-instruction controlled the physical processor by specifying the data sources, data destinations, the ALU operations, and conditional branch actions. The table below shows the structure of the micro-instruction in detail. Note that a micro-instruction controls each component of the processor separately at a low level, so it is very different from a machine instruction. A micro-instruction also provides a degree of parallelism, since it specifies three operations for each step (ALU 1 operation, ALU 2 operation, and a conditional action).

    Format of a 72-bit IOP micro-instruction. From Prototype IOP Functional Description.

    Format of a 72-bit IOP micro-instruction. From Prototype IOP Functional Description.

     

  19. The PROM chips are Intersil IM5624C parts. These are similar to the Signetics 82S131 and Intel 3622 parts. The front side of the page also contains nine chips labeled "D1-6605-2", probably manufactured by Harris; perhaps these are buffers. 

  20. The Prototype Input/Output Processor Function Description lists two pages associated with microcode: "Micro Store (ROM)" (the page that I examined), and "Micro Store Page". I assume that the second page held the 512 words that didn't fit on the first page, along with the circuitry for the microcode control logic and registers. 

  21. Why are the numbers on the PROM chips semi-ordered but also somewhat random? My hypothesis is that the original chips were numbered 1 through 36 in sequence, but when chips needed to be replaced for software patches, each new chip received the next number in sequence, up to 74. 

  22. With flat-pack ICs, an IOP board can hold up to 20 ICs per row, so 100 ICS on a board and 200 ICs on a double-sided page. With the larger DIP packages, the PROM page holds just 45 ICs. Since DIPs are taller (thicker), the page has only a single board. This shows the large density advantage of flat-pack ICs over DIP ICs.

    The density of this page is slightly better because there are a few (15) flat-pack ICs mounted on the back of the PROM board (below). The flat-pack ICs had to be mounted between the rows of DIPs to avoid the pins of the DIP ICs. Because DIPs use through-hole mounting, their pins exit the back side of the board. The large two-pin packages above and below are decoupling capacitors, filtering the power to the ICs.

    Back of the PROM page.

    Back of the PROM page.

    The back side of the board also shows that the printed-circuit board is an inch smaller than the space available; note the gap on the right. Perhaps the circuit board was designed for a standard 8-inch 4 Pi page, but then mounted on the IOP's special 9-inch page. 

  23. The NASA Office of Logic Design web page has a photo of a Space Shuttle board that might be from the IOP, but its source is unknown (I asked). This board is puzzling because it has the same unusual 9" form factor as the IOP pages, but it also has many differences, so it probably came from a different Shuttle system.

    A Space Shuttle board. Note the broken connector; the plastic on these vintage Burndy connections is very often broken. From Space Shuttle Computers and Avionics.

    A Space Shuttle board. Note the broken connector; the plastic on these vintage Burndy connections is very often broken. From Space Shuttle Computers and Avionics.

    The board is a dual MIA interface; it is labeled "ADPTR. INTFC. DUAL MUX", part number "A538A762-02". This part number does not appear in the IOP documentation, and has a different format from IOP part numbers. The circuitry on the board is very similar to the IOP's interface board, with hybrid modules, transformers, and analog components. Physically, the board has the same dimensions, mounting hardware, and 120-pin connector as the IOP boards. However, the board doesn't have the test point connector at the top and the ICs are arranged haphazardly, instead of in uniform rows, so it doesn't look like it was manufactured by IBM. Moreover, the number of ICs is much smaller. On the other hand, it uses the same 54LS395 4-bit shift register chips (labeled 6088913). I would think that this was a prototype board for the IOP's board, except both boards are from 1976, based on the component dates.

    My current hypothesis is that this board was the MIA network interface in a different Space Shuttle component, probably the MDM (Multiplexer/Demultiplexer); the MDM contained a "Serial MIA" board built by Singer-Kearfott. Note that the board has Singer hybrid modules; since Singer-Kearfott invented the MIA network, it makes sense that their modules would be on an interface board. Another possibility is that this board was part of the Shuttle's IMU (Inertial Measurement Unit), which was built by Singer-Kearfott. The IMU communicated with the MDM via a serial I/O line that was very similar to the MIA protocol, but had some differences.

    Singer, by the way, is the same Singer that builds sewing machines. How did they end up making advanced components for the Space Shuttle? (Not to mention nuclear missile guidance systems.) In the 1960s, Singer diversified into defense and computers; in 1968, Singer acquired Kearfott, a defense company that built inertial navigation systems. The Singer-Kearfott SKC-2000 computer was considered for the Space Shuttle, but IBM's AP-101 was selected instead. Singer-Kearfott built the Inertial Measurement Units (IMUs) for the Space Shuttle. In 1987, Singer sold its Kearfott Guidance & Navigation division to the Astronautics Corporation. Kearfott still produces guidance and navigation systems, such as the inertial navigation system for the Global Hawk UAV and the Trident II submarine-launched ballistic missile. After a 1987 takeover and two bankruptcies, Singer is back to just sewing machines, now part of the SVP Worldwide sewing machine company. 

  24. Bonus photo of Peter Kogge working on the I/O Processor:

     

Shape the Future: Call for Submissions Open for WikiConference India 2026

Assamese wikimedians showing the wiki sign: ByUser:Gitartha.bordoloi, Creative Commons Attribution-Share Alike 4.0 International license.

WikiConference India 2026 is now open for Programming submissions!

This conference is a space to create, share, and lead, not only to attend. This year, we are building around the theme Re-imagining the Knowledge Commons, and extending an invitation to you to find out together how we build and care for shared knowledge. We’re looking for proposals that go further than traditional editing, uploading, and coding, trying out new ways of taking part and making knowledge as a community. 

WikiConference India will happen in Kochi, Kerala on 4th- 6th September 2026, in the spirit of Namukku Othukoodam (Let’s meet!)

The Evolving Knowledge Commons

The Knowledge Commons, in its digital form, has come to be defined by the Wikimedia ecosystem. Emerging trends in AI-based search and summarization, content generation, language translation, and new forms of interaction have brought new challenges to the way in which we contribute to the knowledge ecosystem and the way in which consumers unknowingly draw from it. 

To ensure that knowledge remains open, inclusive, safe, and accessible, we cannot simply maintain what exists: we must actively shape ourselves to tackle what comes next and this is at the core of our Programming vision. Our theme aligns closely with broader Wikimedia movement priorities and also the global trends tracked by the Wikimedia Foundation. In a time of declining trust in online information and the rapid rise of AI-generated content; strengthening community-led, human-created knowledge becomes more critical than ever.

Key Questions We Are Exploring

  • How efficiently are we bringing traditional stores of knowledge- from people, libraries, academia, galleries, and museums- into the Knowledge Commons?
  • How do we protect the integrity of our knowledge, keep the Knowledge Commons verifiable, and protect the agency of our contributors?
  • What do all these changes and future possibilities mean for those of us contributing to Wikimedia projects focussing on Indic languages?

Core Focus Themes

We welcome submissions that address our core objectives and align with the following strategic directions:

  • Strategic Roadmap Building: Sessions that help us navigate the shifting digital landscape, the rise of new technologies, along with Sustainability and future pathways for Wikimedia in India and South Asia.
  • Community Leadership and Governance: Sessions that highlight community-led initiatives and how grassroots leadership can sustain open knowledge.
  • New Models of Participation and Contribution: Sessions related to technological enhancement, GLAM, EduWiki, and Knowledge-resource-content partnerships.
  • Knowledge Equity and Inclusion across Languages and Regions: Sessions focused on strengthening Indic-language projects and ensuring underrepresented histories remain at our heart.
  • Technology, Tools, and the Evolving Digital Knowledge Ecosystem: Exploring infrastructure, platform engineering, and the role of AI.

We especially encourage proposals that:

  • Share collective practical experiences and learnings, offering generalizable knowledge.
  • Highlight independent/community-led innovations that are reproducible in wider contexts.
  • Create space for dialogue, skill-building, and co-creation that travels beyond a single community.

If you have received a scholarship to attend, we especially encourage you to propose a session or be a part of one! Whether you are an experienced contributor or a newer community member, your ideas and experiences are essential to building a meaningful and inclusive program.

Submission Tracks & Formats

To help organize our collective program, we invite proposals across a variety of formats. Please review the session formats and suggested durations below to see where your idea fits best:

  • Workshops: 60 – 90 minutes
  • Panel discussions: 45 – 60 minutes
  • Lightning talks: 5 minutes
  • Interactive activity: 30 – 45 minutes
  • Facilitated Discussion / Roundtables: 45 – 60 minutes
  • Demonstrations: 15 – 20 minutes
  • Poster presentations: Posters will be presented at the Conference Venue, allowing one-on-one discussions with participants.
  • Community meetups

Note: Unconference sessions are intended for informal meetings after 5:00 PM (following main conference hours) and will not be listed in the primary program schedule. Even if your idea does not fit perfectly into one of these exact buckets, we still want to hear it!

Perspective from the Movement

Nitesh Gill and Nivas at Wikiconference India 2023: By Biswarup Ganguly, Creative Commons Attribution 3.0 Unported license

“WikiConference India is, at its core, a space for communities to meet, reflect, and reimagine- not just a conference, but a moment where communities see themselves and the impact of their work more clearly. WCI 2023 brought that vision to life by bringing Wikimedians from across India and South Asia together in Hyderabad for the first time since 2016, under the theme ‘Strengthening Bonds.‘ It reminded us that reimagining the knowledge commons is ultimately about people, how we connect, collaborate, and care for the spaces we build together. We hope WCI 2026 continues this journey forward.” – Nitesh and Nivas, Organisers of WikiConference India 2023

“WikiConference India 2026 is shaped by the community, and the program is at its heart. Through this call, we invite contributors to bring their ideas, experiences, and questions to the table and help co-create the conversations that matter most.”

Core Organizing Team, WikiConference India 2026

Key Details & Deadlines

As we prepare to meet in Kochi this September, remember that the “Knowledge Commons” belongs to you. We cannot reimagine it without your voice, your leadership, and your creativity. Submit your proposal today and help shape the roadmap. 

From Application to Notability: My Journey Through the Wiki Afrodemics Mentorship Programme, Cohort 1

By: Umar2z
Ibjaja055, CC BY-SA 4.0 https://creativecommons.org/licenses/by-sa/4.0, via Wikimedia Commons
Ibjaja055, CC BY 4.0 https://creativecommons.org/licenses/by/4.0, via Wikimedia Commons

When I first came across the call for applications for the Wiki Afrodemics Mentorship Programme, I almost scrolled past it. I had seen calls like this before, applied to a few, and heard nothing back. But something about this one felt different maybe it was the clarity of the focus countries, or maybe it was simply timing. I applied anyway, not expecting much, and a few weeks later found my name on the list of Cohort 1 fellows for the Anglophone group.

That single email changed the next three months of my life on Wikimedia.

The Application and the Wait

Applying was straightforward enough: a form, a short statement of interest, a bit about prior contributions. The harder part was the waiting. Mentorship programmes like this one are competitive, and I remember refreshing my inbox more often than I’d like to admit. When the selection list finally came out, and I saw my name among the twenty participants chosen from across the five focus countries, it felt like validation not just of my interest in Wikimedia, but of the small, scattered edits I had been making before anyone was watching.

A Setback I Didn’t See Coming

Not long into the programme, I ran into a wall I genuinely didn’t expect: I was temporarily blocked on Wikimedia. I won’t pretend that moment didn’t sting. For a brief period, I couldn’t edit directly, and I had to sit with the uncomfortable feeling of being sidelined from the very project I had just been selected to contribute more to.

But this is where the structure of the programme and the patience of the mentors made all the difference. Instead of treating the block as a dead end, I learned to treat it as a detour. I kept writing. I drafted articles offline and in my sandbox, and submitted them for review by experienced Wikimedians who could push them through on my behalf or guide me on how to get back in good standing. It taught me something I didn’t expect to learn from a mentorship programme: that contributing to Wikimedia isn’t only about the edit button. It’s about the work itself: the research, the sourcing, the drafting, and there is always a way to keep that work moving, even when the front door is temporarily closed.

Mentors Who Knew How to Teach, Not Just Tell

If there’s one thing that stood out across the three months, it’s the quality of the mentorship itself. It’s one thing to know Wikimedia inside and out; it’s another to be able to teach it well. Our mentors managed both.

The Wikidata sessions, in particular, stretched my thinking. Wikidata isn’t always intuitive when you’re coming from a Wikipedia-first mindset, but the structured, month-long training broke it down in a way that made the structured data side of the movement click for me. Moving between Wikipedia, Wikidata, Wikimedia diff and Wikimedia Commons over the course of the programme gave me a much fuller picture of how the projects connect, how an image uploaded to Commons, a claim modeled on Wikidata, and an article written on Wikipedia all reinforce each other.

Ibjaja055, CC0, via Wikimedia Commons
Ibjaja055, CC0, via Wikimedia Commons
Ibjaja055, CC0, via Wikimedia Commons
Ibjaja055, CC0, via Wikimedia Commons

The Lesson I’m Taking With Me: Notability

Of everything covered across the cohort – content gaps, gender representation, cross-regional collaboration – the single most valuable thing I walked away with is a real, working understanding of Notability.

Before this programme, notability was a word I associated with rejection, articles I have seen tagged for deletion, drafts that sat untouched because I wasn’t sure they would survive scrutiny. Now I understand it as a framework rather than a gatekeeping obstacle. Knowing how to evaluate whether a subject meets Wikipedia’s notability guidelines, and how that maps differently onto Wikidata’s notability expectations, has changed how I choose what to write about in the first place. I no longer draft and hope. I check first, source deliberately, and build a stronger case for inclusion from the very first sentence.

That shift alone from guessing to evaluating is worth everything else I gained from this cohort.

Looking Ahead

As Cohort 1 closes out for 2026, I am left with a mix of gratitude and momentum. Gratitude for mentors who gave their time generously over three months of training, listening, and patient correction. Momentum because I now have both the skills and the confidence to keep contributing, blocks, setbacks, and all.

I am already looking forward to the 2027 cohort, not as a participant this time, perhaps, but as someone who might be in a position to give back the same kind of guidance I received.

Ibjaja055, CC0, via Wikimedia Commons

Digital policy directly impacts how we live together as a society

By: A2namzg
Claudia Garad

For over ten years, Wikimedia communities and associations in Europe have been working together to improve the legal framework for free knowledge in Europe – advocacy at the European level and participation in shaping policy at the national level go hand in hand. Copyright reform, the AI ​​Act, and the Digital Services Act (DSA) are just a few examples that have occupied us in recent years. 

Claudia Garád, Wikimedia Österreich’s Executive Director, is an active supporter of the European idea within the Wikiverse. Since 2022, Claudia has served as the volunteer President of Wikimedia Europe – a European umbrella organisation in Brussels, shaped and managed by over 30 European Wikimedia associations, and acting as a strong voice for public-interest-oriented internet policy at the European level. 

Claudia, what exactly does Wikimedia Europe do? 

Our aim is to be a kind of network hub in the international Wikiverse, facilitating effective exchange and cooperation between Wikimedia organizations. Decentralized cooperation is a major strength of our volunteer communities. In Europe, we demonstrate that Wikimedia organizations also have this way of working in their DNA. Especially in these politically and socially challenging times, when civic engagement is increasingly restricted everywhere, we can only survive and exert real influence as a collective. 

Wikimedia Europe is therefore a shared platform for collaborative work on European legislative processes that impact Wikimedia projects. It also supports collaborative fundraising activities for joint projects and skills development and transfer. This happens particularly in the areas of advocacy and fundraising, but also beyond. 

Why is international digital policy important for an affiliate such as Wikimedia Austria? 

First and foremost, it’s about the digital infrastructure on which everything is built: For Wikipedia, other Wikimedia projects, or our ÖsterreichWiki to function at all, we need an open, free, and globally functioning internet. Without interoperability and open standards, these projects would not be possible. European and international digital policy, therefore, determines very concretely whether free knowledge remains accessible to everyone and under what conditions our communities can operate. 

Furthermore, national legislation in this area is largely based on European regulations, which are usually designed with large, commercial platforms and social media in mind. Non-commercial, public service digital projects are often not given a voice and cannot lobby to the same extent as multi-billion-dollar corporations. Through Wikimedia Europe, we have the opportunity to pool and amplify the activities and resources of individual members and thus counteract this imbalance. 

Last but not least, it’s also about transparent political decision-making processes: Digital policy doesn’t just affect technical or economic developments, but directly impacts how we live together as a society. After all, we use digital platforms for consumption and communication, we form our opinions online, we work with digital tools, and so on. Therefore, we advocate for the structural integration of organizations representing civil society into decision-making processes in Austria and Europe: Politics, business, science, and civil society should work together transparently on solutions. 

What is your role at Wikimedia Europe? 

During the founding phase, I initially served as president on the interim board to prepare the organization for its spin-off, and then in 2025 I was elected to the first “official” board in the same role. My responsibilities include representing the association externally—although this has now been largely delegated to the managing director, Anna Mazgal, in day-to-day operations. In addition, I handle governance and HR matters, as well as risk management within the association, and act as a liaison to the global Wikimedia movement beyond Europe. Last year, together with the board, employees and a project group of representatives from the member organizations, we also developed our first integrated multi-year strategy. 

What does Wikimedia Europe mean to you personally? 

Jacques Delors, former President of the European Commission and architect of European integration, used to say: “Never choose between being an optimist or a pessimist. The only choice you can make is to be an activist.” 

My volunteer role as President of Wikimedia Europe gives me the opportunity to be an activist in two projects that have been formative for me personally and represent, in my opinion, the most wonderful experiments in human history: Wikipedia and the European Union. Both are examples of how people can achieve the previously unimaginable when they consistently prioritize trust and cooperation. This gives me courage, confidence, and a sense of self-efficacy, even on dark days when I feel the world is increasingly falling apart. 

To learn more about WMEU visit our page on Meta-wiki

Check the WMEU public policy blog to follow the latest developments in our public policy work in Europe

Peter Zlabinger is a Communications Advisor at Wikimedia Österreich and his job is to explain the ins and outs of the Wikiverse to our various audiences and stakeholder groups. In EU advocacy two complex universes meet – the policy making of the European Union and the Wikimedia Movement. Wikimedia Österreich has been an integral part of Wikimedia advocacy in Europe. In this interview Peter explores what makes advocacy so exciting but also important from an Austrain but also global movement perspective.

From Rock Art to Digital Knowledge: Wikimedia Foundation CEO Visits South Africa

The Wikimedia movement’s commitment to preserving and sharing knowledge took on a uniquely South African flavour last month as newly appointed Wikimedia Foundation CEO Bernadette Meehan embarked on one of her first major international visits since assuming the role.

Accompanying her was Bobby Shabangu, a respected South African open-knowledge advocate and the first African ever elected to the Wikimedia Foundation Board of Trustees. Together, they spent time with Wikimedia South Africa volunteers, partners, and community members, exploring how local initiatives are helping ensure African languages, histories, and cultures are represented online for generations to come.

Cape Town: Where Ancient Stories Meet the Digital Future

The South African tour began in Cape Town with a visit to the Iziko Museum, where participants experienced a guided tour that perfectly reflected the Wikimedia movement’s belief that knowledge is created, shared, and preserved by people.

Among the highlights was a presentation of some of South Africa’s oldest rock art. Created over centuries by multiple generations, these artworks offered a striking parallel to the way Wikipedia itself is built today — collaboratively, incrementally, and collectively. Just as countless hands contributed to preserving stories on stone, volunteers across the world continue to build humanity’s largest collection of freely accessible knowledge.

Conversations with Cape Town-based Wikimedia South Africa board members and partners took visitors further “down the WikiRabbit Hole,” exploring innovative language preservation projects currently underway. Particular interest was shown in efforts supporting the incubation of Afrikaaps, a Cape Town dialect of Afrikaans, as well as initiatives enabling KhoeKhoe language speakers to contribute directly online through the development of specialised keyboards and editing tools.

Johannesburg: Showcasing Local Innovation

The visit continued in Johannesburg, where Wikimedia South Africa members presented a range of projects addressing the digital knowledge gap facing African languages and communities.

Many South African languages have rich oral traditions but limited representation in written and digital archives. Community members demonstrated how Wikimedia projects are helping address this challenge through locally led initiatives.

Decolonising Language Through Knowledge Creation

Editors showcased work on smaller and underrepresented language editions of Wikipedia, including Siswati, Tshivenda, Sepulana, and Xitsonga. Through the creation of high-quality educational content in local languages, volunteers are helping ensure that knowledge is accessible to communities in the languages they speak and understand.

Jo’burgpediA and Community Partnerships

The chapter also highlighted collaborations with universities, libraries, museums, and cultural institutions. Through initiatives such as Jo’burgpediA, students and community members are trained to document local history, ensuring that communities can tell their own stories rather than having them told by others.

Preserving Oral Knowledge

Discussions also focused on innovative approaches to documenting verified oral histories and indigenous knowledge. These efforts challenge traditional archival models that have often excluded African perspectives and lived experiences, creating new pathways for preserving cultural heritage within Wikimedia platforms.

A Global Mission Powered by Local Communities

Addressing chapter members, Bernadette Meehan emphasised the Foundation’s commitment to community-led knowledge creation and governance:

“Wikimedia’s global mission relies entirely on the local communities who understand the nuances of their own culture. Listening to the brilliant initiatives run by Wikimedia South Africa proves that the future of free knowledge is inherently multilingual and diverse.”

Her remarks recognised the vital role that local volunteers play in ensuring the internet reflects the richness and diversity of the communities it serves.

For Bobby Shabangu, the visit was also a reminder of how far the African Wikimedia movement has come.

Reflecting on his own journey, which began with editing Siswati Wikipedia because, as he puts it, “if I didn’t edit it, no one would,” Shabangu highlighted the growing influence of African contributors within global internet governance and knowledge-sharing spaces:

“We are moving past the era where Africa is merely a consumer of global knowledge. Through the hard work of our chapter members, we are ensuring that our grandfathers’ stories, our traditional customs, and our beautifully diverse languages are permanently archived for the next generation.”

Looking Ahead

The visit reaffirmed the importance of community-driven knowledge creation and the growing role South Africa is playing in shaping the future of free knowledge globally.

From ancient rock art in Cape Town to digital language preservation projects in Johannesburg, the message was clear: preserving knowledge is not only about recording the past — it is about ensuring that future generations can access, contribute to, and share their own stories.

As Wikimedia South Africa continues to champion linguistic diversity, cultural heritage, and open knowledge, partnerships between local communities and the Wikimedia Foundation will help ensure that South Africa’s rich tapestry of languages and cultures remains visible, accessible, and thriving online for decades to come.

Wiki Indaba 2026: Passing the Baobab

Joining the Wikimedia South Africa chapter were members of the core organising team for Wiki Indaba 2026 from Côte d’Ivoire, including Emmanuel Gueh and Donotein. Their participation provided an opportunity not only to engage in discussions with the South African community but also to share an update on preparations for Wiki Indaba 2026, Africa’s premier gathering of Wikimedians.

A special highlight of the visit was the ceremonial handover of the Wiki Indaba Baobab Tree statue to the new organising team. The Baobab, often referred to as the “Tree of Life” in Africa, symbolises wisdom, resilience, community, and the sharing of knowledge—values that closely align with the spirit of Wiki Indaba and the Wikimedia movement.

Wikimedia South Africa leadership proudly passed the statue to the Côte d’Ivoire organising team, marking the beginning of what is hoped will become a lasting Wiki Indaba tradition.

As the creator of Wiki Indaba, Wikimedia South Africa Board Chair and long-time Wikimedian Dumisani Ndubane reflected on the significance of the moment. He expressed his pride in inaugurating this new tradition and shared his hope that, in years to come, the Baobab statue will continue its journey across the continent before eventually returning to South Africa when Wikimedia South Africa once again hosts Wiki Indaba.

The handover symbolised more than a transfer of responsibility—it represented the continued growth of the African Wikimedia movement, the sharing of leadership across communities, and a collective commitment to ensuring that Africa’s knowledge, languages, cultures, and stories remain visible and accessible to future generations.

I built the convertor from Wikipedia dumps to man (roff) format – for offline reading in a terminal

I love Wikipedia, command-line interface (cli) and respect offline stuff. Finally I did it – Rust software (for peformance), code generated by llm gpt-5.5 xhigh.

How to use: download the Wikipedia dump, and

cargo run --release -- path/to/dump.xml -o path/to/output_directory

Yes for offline reading we have Kiwix – but terminal is also good. For you batery as well. Or when you system in under heavy load so your CPU/RAM are limited. Man format mean that you can read Wikipedia through SSH as well, and from very cheap devices.

This is not perfect – we should add special handlers for different templates. But it already mostly works and valuable for the comunity. Try it, ping if you want to improe something. You are welcome to pack this software for you linux distribution – I packed only for Gentoo.

Wikipedia in man screenshot

See the project repository, with instructions, at https://gitlab.com/vitaly_zdanevich_wikimedia/wiki2man_on_rust

From Learning to Contributing: My Wikimedia Mentorship Journey

This is a testimonial graphic design for the On-Wiki Skills Mentorship Program Cohort 2 (2026)

Before joining the On‑Wiki Skills mentorship program, my engagement with Wikimedia projects was mostly as a reader. I was curious about open knowledge but unfamiliar with how data works or how contributors actively build and maintain Wikimedia projects.

Through the mentorship, I was introduced to Wikimedia as a collaborative ecosystem that offers structured ways of organizing and connecting knowledge. The learning curve was initially challenging especially understanding items, properties, statements, and references but with guidance from mentors and hands‑on practice, these concepts gradually became clearer.

Skill Acquired


During the training, I developed skills in sourcing and referencing information and understanding structured data. These experiences changed how I view knowledge creation and reinforced the idea that learning continues beyond formal lessons.

Some of my image uploads

  • Administration Block from Ghana Standards Authority
  • Pesticides Residue Block from Ghana Standards Authority
  • Fish Department from Ghana Standards Authority

Participating in practical exercises and guided contributions helped build my confidence and showed me that I could engage meaningfully with a global open‑knowledge communication.

My Gratitude

I am grateful to the Wikimedia community, mentors, and program organizers for their support and encouragement. Completing the mentorship marked a new beginning for me, and I look forward to continuing to explore and contribute to open knowledge

Documenting Nigeria’s Rites and Rituals through Wiki Loves Africa 2026 for Creatives in Nigeria

By: Meritkosy

For years, rites and rituals have shaped the identity of communities across Nigeria. From naming ceremonies and traditional weddings to religious observances and everyday customs, these practices preserve values, beliefs, and collective memories. Yet many of these cultural expressions remain underrepresented on Wikimedia projects.

 Wiki Loves Africa 2026 for Creatives in selected Nigerian universities broughtt together photographers and storytellers to document and share Nigeria’s rich cultural heritage under the theme “Rites and Rituals.” The campaign was held from February to April 2026.

Creating awareness and recruiting creatives

Between 1 and 28 February 2026, the campaign focused on awareness and participant recruitment. Outreach activities targeted creatives across six institutions and communities, while social media campaigns and sponsored advertisements on Facebook and Instagram helped expand participation.

These efforts introduced creatives to Wikimedia Commons and encouraged them to contribute freely licensed media that celebrate Nigeria’s diverse traditions.

Flyer for Awareness

Launching the campaign and introducing Participants to Wikimedia Commons

The campaign officially kicked off with a virtual launch held on 4 March 2026. The session introduced participants to the Wiki Loves Africa 2026 theme, “Rites and Rituals,” and provided an orientation on contributing to Wikimedia Commons.

Participants learned about:

  • Wikimedia Commons and its role in preserving open knowledge;
  • Creating Wikimedia accounts;
  • Understanding free licenses;
  • Best practices for successful contest submissions.

The virtual launch provided a foundation for participants, which some were first-time contributors to Wikimedia projects.

Training sessions

Throughout March and April, participants joined several training sessions organized by the international Wiki Loves Africa team. These sessions equipped creatives with technical and ethical skills required for documenting cultural heritage responsibly.

Some of the sessions included:

  • The Wiki Loves Africa International Virtual Launch on 5 March 2026;
  • A photo essay writing workshop on 11 March 2026;
  • Audio creation and sound design training for Wikimedia Commons on 19 March 2026;
  • Ethical storytelling training on 27 March 2026;
  • “Documenting our Rites and Rituals with Respect and Integrity” on 28 March 2026;
  • Pattypan mass upload training on 4 April 2026;
  • Uploading photographs, videos, and audio files to Wikimedia Commons on 10 April 2026.

These learning opportunities strengthened participants’ understanding of visual storytelling, copyright, ethical documentation, and the technical processes involved in contributing multimedia content to Wikimedia Commons.

Taking the campaign to communities through photowalks

Beyond virtual engagement, the campaign emphasized practical documentation through physical meetups and photowalks.

Creatives in Rivers State held their meetup and photowalk between 3–5 April 2026, where participants explored their communities and documented the Easter Rites.

On 11 April 2026, creatives in Abuja gathered for their meetup and photowalk, providing opportunities for collaboration, peer learning, and hands-on experience in capturing images related to the theme. Gombe creatives hosted their physical meet-up on 25th April 2026, and so did Edo State and Oyo State.

Physical meet-up of creatives in Abuja
Physical meet-up of creatives in Abuja
During physical meet-up

Every photograph, audio recording, and video uploaded to Wikimedia Commons contributes to a growing repository of freely accessible knowledge, ensuring that Nigeria’s rites and rituals are visible to people around the world and preserved for future generations.

As participants continue contributing beyond the campaign, Wiki Loves Africa remains a reminder that documenting culture is not only about preserving the past. It is also about ensuring that African stories are represented and shared openly with the world.

Submissions can be explored through the event page , the dashboard or through the category

Commons Quality Images, 20 years of celebrating Wikimedians taking those photographs we need

By: Gnangarra

Commons Quality Images (COM:QI) was created during June 2006. The purpose of COM:QI was to identify high quality photographs taken by Wikimedians and made freely available to the community at large through Commons. As of 14 June 2026 over 450,000 media files, mostly photographs, have been recognised through this process. Alongside photographs, QI has Microscopic images, animated GIF’s, and graphical works including our own QI seal (pictured).

Green wax seal
Commons Quality Image seal

From the Beginning

In September 2005 Commons became live, just 6 months later I was encouraged over by User:Pfctdayelise who had been there from the first month. In June 2006 User:Pfctdayelise was searching to put together collections of images created by Commons users as rotating background images, or calendars to help promote the project. I tried to help, we could not even find a common subject matter consisting of reasonably sized and quality printable images. Wikimedia Commons had some great photos, some of which had been featured, but a substantial portion were images scraped from other sites like NASA and Flickr.

That started me thinking about how we raise the profile of community uploaded images and encourage efforts to improve them. Having already contributed to some Good Articles over on en.Wikipedia I thought we could create a similar project but for photographs.

Creating the project I saw some key necessities: it needed to be efficient with clear standards, while not taking weeks to decide the outcome. Most importantly it needed to focus only on images that had been taken by members of the Commons Community and that once recognised it could not be taken away. From there working with User:Wikimol and others we built the project. We also created image guidelines which would enable consistent reviewing of nominations.

The image guidelines started out as;

  1. Photographic images
    1. Low resolution / thumbnail. Photographic QP has to have at least 1.92 megapixels (=1600×1200, for example)
    2. JPEG problems. Too much compressed, too low JPEG quality settings in camera / when saving. Visible jpeg artifacts. => Use better quality settings (e.g. set JPEG “superfine”, shoot RAW, save in photoshop with max. quality)
    3. Noise problems, too much noise. Be it chroma noise, luminance noise, visible grain, scratches in scans… QP should not have distracting amount of noise when viewed in 100%
    4. Bad exposure. Overexposure, blown out highlights, underexposure, shadows details replaced by jpeg maps… In incorrectly exposed images, significant details in a significant part are lost.
    5. Color problems. Bad white balance. Distracting (typically purple) hazing at 100%. Color aberration. QP must have reasonable colors (which does not necessarily mean natural colors).
    6. Improper or undefined focus, insufficient depth of field. QP should have clearly defined focus, e.g. main subject in focus, foreground and background out of focus. Or the whole scene in focus. Counterexample – main subject blurry, foreground even more blurry, focus is somewhere between main subject and background. DOF could be low on purpose.
    7. Blur. Images blurred just because of shaking hand or subject moving too fast. Motion blur in QP has to have purpose.
    8. Poor lighting. Including: distracting reflections (usual problem with built-in flash), unintended vignetting, distracting harsh shadows. Generally bad lighting makes scenes with space look flat.
    9. Overfiltered. There are so many PS/Gimp filters. Rarely a better image is created just by applying more and more filters…
    10. Bad or nonexistent composition, unclear or nonexistent subject. QP should have subject and composition of the image should support depiction of that subject, not distract from it.
    11. Bad perspective, tilt, and other distortions. An eye (or, more precisely, a brain) is a sensitive detector capable of spotting even a small tilt … falling trees, churches, inclined water surfaces,… Images of architecture should usually be rectilinear and without too much perspective distortion.
  2. Stitched images, panoramas.
    1. Panoramatic QP has to have a height of 800px min.
    2. Stitching problems. Stitched images should be without artifacts, colors and lightness should be the same across the image.

All of these may look familiar to many of you as they are still present as Commons:Image guidelines, the very same guidelines that every WikiLoves or similar competition uses as their rules. 

Now we had a process with guides, we approached User:LadyofHats who was known for the drawings she was uploading at that time. LOH was asked to create a seal which we could use to help identify successful photos, that seal is the one we continue to use. Some time during this the QI seal itself got recognised as a Quality Image.

In those early months every image was reviewed and the pages manually updated, a time consuming process that would occasionally be interspersed with edit conflicts. The community embraced QI and a new project called Valued Image emerged for recognising sets of image rather than individual images. It was the efforts of User:Dschwen, who had also taken on maintenance tasks, decided to create a bot that would do most of the work for us. Mike Peel continues to maintain the QICbot and has been invaluable at keeping it working, as of 16 June 2026 QICbot has performed 922,000 edits looking after QI tasks.

This photograph of a Western Gull (Larus occidentalis) was the first Quality image to be promoted to Featured Picture status by User:Dschwen

QI grew fast once the QICbot came on board. I stepped into the background watching the community grow QI. Over time many contributors have made QI. Like User:Poco a poco who presented at Wikimania in London on how QI helped him improve his contributions. For the curious there is an opt-in unaudited list of QI by photographers. Whether it’s one successful image or 20,000 of them they all make QI what it is. 

The future

I think some of QI’s potential still hasn’t been realised. I saw it as a historical record of the growth of photography, of something researchers could look back over and see how it has matured. Perhaps more could be done to integrate QI images in content on other projects as they are recognised as our better works. Maybe a bot could triage nominations to identify the more regular issues that cause images to be rejected, though the human touch should always remain the final judge. 

In the future, QICbot will reach 1 million edits, QI will reach 500,000 very soon. QI carries within itself a lot of untapped potential. 

Could it be time to consider creating a sister project for Quality Videos? I know one thing: there will always be enough high quality images created by Wikimedians to make a calendar for any subject. The challenge will be in choosing just 12.

I would must take a moment to acknowledge everyone who has already helped Commons Quality Images along it’s 20 year journey there has been so many as well as everyone who joins the efforts in future. It’ll always be nice to reflect and be able to say I was able to open the door to something so special. The future of Quality Images is now a journey the Wikimedia Commons Community will decide. 

QI’s prosperity comes from the collective effort of everyone! Perhaps this will encourage you to join the QI club.

Team Challenges: New Perspectives and Innovation at Wikimania 2026

For 25 years, the Wikimedia movement has been constantly innovating to promote the values of free and universal sharing of all human knowledge for the benefit of humanity. In 2026, as we face the challenges of connectivity, misinformation, and the rise of AI, our ability to innovate collectively is more essential than ever.

To address these developments, we are introducing a new program focusing on Team Challenges in the lead-up to Wikimania. This year, we are inviting experts and newcomers to the Wikimedia community to join forces with participants in our Wikimania Hackathon. The idea is not to replace the hackathon format, but to offer a different approach. We want to bring together diverse perspectives and skills by pairing Wikimedians from all walks of life (not just developers!) with professionals from other fields. The goal is to share our best practices and design digital tools that are sustainable, inclusive, and accessible to all.

Next week, we are kicking off our first online orientation sessions training newcomers on wiki tools. In the coming weeks, they will form teams with Wikimedia community members to undertake one of the 2026 challenges. Stay tuned for further updates as the teams work together toward the Team Challenges showcase concurrent with Wikimania Paris. 

We can’t wait to see what the cross-pollination between different disciplines, perspectives, and skill sets will bring. Good luck to all!

10 challenges for maximum impact

The goal is simple: To turn the community’s needs into concrete solutions. We have selected 10 major technical challenges that address the complex issues encountered on a daily basis across Wikimedia projects. 

These challenges are based on the three pillars of Wikimania 2026:

  • FREEDOM: Ensuring the interoperability of free software to facilitate its reuse and maintenance, and to guarantee its long-term viability.
  • EQUITY: Promoting multilingualism. Designing tools capable of overcoming language barriers to bridge gaps in sources and contributions on a global scale.
  • RELIABILITY: Developing AI-based detection tools to assist volunteer moderators in their work of verifying and protecting information.
Discover the 10 Challenges
Looking at a blue puzzle piece through a magnifying glassFix the sources Automatically detect outdated, retracted or broken sources in articles, and facilitate their updating or archiving.

Child and older man sitting at a table doing a puzzle together
Welcoming newcomers Reinvent how we onboard newcomers through chatbots, guided micro-edits, and gamification lower the barrier to entry on wikis.
Birthday cake to celebrate the 25th anniversary of WikipediaUpdate the obsolete Identify and flag outdated articles, obsolete illustrations and screenshots, and outdated graphics to: encourage the community to update them.
Wikimania Paris lightbulb character with a megaphoneBoost multimedia experience Enhance the audiovisual experience on wikis: playlists on Commons, a modernised video player, annotation of audio segments, and games focused on accents and pronunciations.
Statue of a thinker with a pigeon on his headDeciphering biases Develop tools to detect gender, geographic, and cultural biases in Wikimedia content to: encourage the community to correct them.
Wikimania Paris lightbulb character kicking a puzzle globe footballGamify knowledge Create serious games and playful experiences using Wikimedia data to: learn, contribute, and explore in new ways.
French waiter with a serving tray of books and knowledgeStream data with Wikidata Take advantage of Wikidata content to: generate lists, tables or up-to-date visualisations, semi-automatically, for other projects in the Wikimedia galaxy or further.
Laptop with Wikimania Paris mascot Luce the lightbulb visible on the screenThe editor of the future Improve the editing experience: search & replace in the source editor, a standardized citation toolbar, a better visual table editor, and text alignment in VisualEditor to modernise and streamline contributions.
Explore knowledge Design and experiment new ways of sharing knowledge, in interactive and visual ways…for the modern web!
Connect multilingual knowledge Facilitate translation, the reuse of content across Wikimedia projects, and collaboration with other Open Source projects so that knowledge can flow freely in all languages.

For more information on the Team Challenges and how to participate, please visit the Team Challenges section of the Wikimania wiki. Please note that registration for the Team Challenges does not grant access to the rest of Wikimania.

Tech News 2026 – Issue 26

Latest tech news from the Wikimedia technical community. Please tell other users about these changes. Not all changes will affect you. Translations are available: 日本語, Bahasa Indonesia, Bahasa Melayu, Deutsch, Tiếng Việt, español, français, polski, português, português do Brasil, čeština, русский, українська, עברית, العربية, 中文

Weekly highlight

  • Growth features are now available at Wikidata. This update enables access to Mentorship (if configured), Impact module, the Help Panel, and a simplified Newcomer Homepage (without Suggested Edits). Wikidata administrators are still configuring the features through Community Configuration.

Updates for editors

  • The special page Special:RangeCalculator has been created. It allows users to find an IP range without needing to rely on external tools. Until now, this tool was only available to CheckUsers. [1]
  • Sub-referencing is a new MediaWiki feature that allows editors to reuse references with different details. It will be deployed to most small and medium-sized Wikipedia language versions on June 23. The FAQ lists possible actions to take on your wiki to support the deployment. Check the rollout plan for the next deployment steps. [2]
  • Starting next week, users will get a notification when they are blocked or unblocked from editing, or if this block changes. [3]
  • Recurrent item View all 32 community-submitted tasks that were resolved last week.

Updates for technical contributors

  • Starting next week, abuse filters that are set to “require CAPTCHA verification” will begin to also affect users with the skipcaptcha right, which includes most autoconfirmed users. Bots are exempted. This change only affects edits that trigger an abuse filter. The skipcaptcha right will continue to exempt users from having to solve CAPTCHAs in the ordinary course of using the wikis. [4]
  • Reference documentation for the Lift Wing API has moved from the API Portal to the interactive REST Sandbox.
  • The API Portal wiki is now closed. For API documentation, see Wikimedia APIs on mediawiki.org. All API Portal wiki URLs (https://api.wikimedia.org/wiki/) will redirect to the mediawiki.org page starting June 22. [5]
  • Recurrent item Detailed code updates later this week: MediaWiki

Meetings and events

Tech news prepared by Tech News writers and posted by bot • Contribute • Translate • Get help • Give feedback • Subscribe or unsubscribe.

from hookswitch to grave

Through decades of consolidation, reorganization, and divestiture, AT&T left a famously complicated corporate history. One of the greatest enterprises in American history, arguably the greatest enterprise, AT&T has often rivaled the federal government in the size of its budget and workforce. One of the reasons, as we well know today, was monopolization and its close relative vertical integration. AT&T was the telephone system, or at least aspired to be, and for decades the meaning of "Universal Service" was that the service was designed, built, and operated by AT&T—universally.

While AT&T's tangled origins are fertile ground for the historian, they also obscure many of the early stories of telephone history. Much of the work of the early independent telephone industry has been lost in the voluminous achievements of AT&T. Even very basic facts become obscure. For example, who invented the telephone? Well, we all know the answer: Alexander Graham Bell. We have mostly forgotten that, at the time, this was a hotly contested question. One of the most prominent alternate claimants to the title was a man named Elisha Gray, today immortalized as the "Gray" in electrical distributor "Graybar," but better known in his time as an inventor of telegraph and telephone equipment. Gray contracted prototyping of some of his inventions to an upstart manufacturer and de facto Western Union spinoff, founded by Enos M. Barton (the "bar" in Graybar) and George Shawk. Impressed by Barton's operation, and at odds with Shawk on its future direction, Gray put together the money to buy out Shawk and became half-owner of the company that would reincorporate, in 1872, as Western Electric (WE).

It is ironic, of course, that a man who might fairly be called one of the top enemies of Bell helped to found the company that would become one of the most important parts of the Bell System. It's not a coincidence: Gray's involvement in WE included plans to manufacture his own telephone design, for which he had filed a provisional patent. Like many of the late 20th century's telephone inventors, Gray's greatest challenge in commercializing his invention was not technical but legal. His provisional patent on a telephone transmitter, substantially similar to the one invented by Bell and possibly older, led Western Union to take take part ownership in WE to advance their own plan to compete with AT&T as a telephone company. That set off a protracted legal battle, whose end result included the termination of Gray's patent claim and Western Union's abandonment of telephony.

AT&T was not the kind of company to leave things to chance, though, and least of all when it came to competition. In 1881, AT&T acquired WE. From that point on, WE was no longer a competitor, it was a core part of the Bell System: the manufacturing and supply arm of AT&T. A few decades later, WE had become the primary maker, and often sole supplier, of every piece of equipment used in the Bell telephone network. Everything from telephones to cables to central office switches were made at WE's various works. What few components WE didn't make, it sourced, through an expansive purchasing arm that negotiated orders on the behalf of the entire AT&T family. In 1925, WE had become so dedicated to the Bell System that its remaining non-telephone business, mostly local distributorships, was spun out into a separate company (Graybar). From that point forward, the Bell System was not only WE's sole shareholder but its sole customer as well.

As part of the 1925 reorganization, WE's research and development arm became a new organization, jointly owned by WE and its patron AT&T: Bell Laboratories. This new organization consolidated AT&T's expanding basic science efforts with WE's manufacturing expertise, setting the stage for decades of equipment that was conceived, designed, manufactured, and used within the AT&T empire. Bell operating companies got everything they needed, from tools to the telephones themselves, via requisition to their local WE supply warehouse. Such were the needs of the growing telephone system that WE started manufacturing telephone cable in 1925, quickly became the world's largest manufacturer of wire and cable, and likely held that title continuously until the turn-down of much of its manufacturing capacity in the late 1970s. AT&T was the nation's largest private employer for much of this period, and WE accounted for about 1/6th of that workforce.

Western Electric cable reel

Until the Carterfone decision and, for the most part, until the divestiture of AT&T in 1984, telephones were born at Western Electric. All of the phones leased by Bell Operating Companies, ranging from the classic WE 500 to explosion-proof phones for coal mine applications, were made at WE facilities like the Indianapolis Works. There were nearly 10,000 employees there, making 35,000 phones a day—and Indianapolis was not remarkable. It was just one plant of many.

AT&T built its empire through innovation, but also through domination. The acquisition of WE was one of its biggest steps towards complete integration, a goal that WE would pursue through the middle of the 20th century. For example, when the Morton Salt empire indirectly led to the Teletype Corporation and the development of commercial teletypewriter networks, they bought it. Teletype was a WE company from 1930 to its end.

Telephones were not only born at WE; they went there to die. During the 1920s, telephones were expensive instruments that required regular maintenance. Besides the commercial advantage, which would become more significant in later years, this aspect of telephones encouraged a full service lease model. Customers leased their phones from their telephone company in part because (prior to Carterfone) they had to, in part because the arrangement made the telephone company responsible for the phone's care. At the same time, Bell Operating Companies carefully controlled their expenses by reusing equipment as much as possible.

So, when a customer signed up for telephone service, they were issued a phone. When they canceled service, or had trouble with the phone, or quite simply wanted a phone that was a different color (or an upgrade to a Trimline or a Princess), the telephone company took the phone back. It would join hundreds of other phones on a trip to the nearest WE Service Center. The same truck would likely make the return journey loaded with phones ready for customers: the service center refurbished them.

Millions of telephones come back to Bell System service centers each year, many with their housings, handsets, and other molded plastic components bruised and battered. Some can be put back in shape by buffing, solvent polishing, or painting. Others wind up in piles. (Q1)

The scale of WE's phone refurbishing program was remarkable. Huge workshops of WE employees inspected, cleaned, repaired, and tested each phone. Refurbished units visited a test desk for a thorough electrical checkout before they received the service center's stamp or label that they had been remanufactured for use. In the Mountain States and west, WE service centers were found in Denver, Phoenix, Los Angeles, San Francisco, and Seattle. By the 1960s, Portland and Salt Lake City had joined.

Of course, despite the best efforts of all of WE's horses and all of WE's men, not all telephones can be put back together again. Much of the equipment returned to WE could not be satisfactorily refurbished. Besides, it wasn't just phones that telephone companies returned to WE, it was everything. Upgrading a crossbar exchange to an ESS? The ESS came from Western Electric, and the crossbar exchange went back to them. WE supplied telephone poles to the operating companies, and at the end of their life it took them back.

Western Electric has manufactured millions of telephones, millions of miles of wire and cable, tens of thousands of manual and dial switching units, and the thousand-and-one other kinds of apparatus that go into the plant of the Bell System. It has purchased from thousands of other manufacturers the great variety of supplies that are used by the Bell System. (Q2)

The majority of that output—at least what wasn't still in service during WE's decline—went back to WE for disposal, as well. That included the wire: from simple drop wires to heavy multipair cables, old wiring was routinely cut into sections and shipped back to WE—specifically, to the WE Salvage Works on Staten Island.

In 1883, as the component elements of a telephone industry were swirling around New York and accreting by gravity into the shape of the Bell System, Benjamin Lowenstein arrived from Germany. Settling in New York City, he took up a business that he must have learned back in Europe: metal refining. Within a year of his arrival, the B. Lowenstein & Bro. company was smelting scrap metal from a shop in Manhattan (the brother, Moses Lowenstein, was a constant second fiddle in Benjamin's ventures until he sold his share and retired to go his own way in 1900).

Lowenstein had a way of maneuvering his metals businesses into the path of technological progress. His first such success was lead, or rather an alloy of lead with tin and antimony. This specialized alloy was eutectic, meaning that it melted and solidified at a single, well-defined temperature, and a low one at that. These were exactly the requirements for feeding the newly-invented Merganthaler hot-metal typesetting machines, later known as Linotype—much as the metal came to be known as Linotype alloy. By 1890, B. Lowenstein & Bro. was the major supplier of feedstock for hot-metal typesetting in the US.

Linotype metal brought in a lot of money, enough that Lowenstein looked to expand. Manhattan was already dense enough that it was hard to find a site for a large industrial operation. Instead, Lowenstein found land in the southern end of Staten Island, near the neighborhood of Tottenville. There, he founded the Tottenville Copper Company. Tottenville Copper grew quickly, well positioned for the new demand for copper brought about by the electrical revolution. During the 1900s, Lowenstein rebranded B. Lowenstein & Bro. as the Nassau Smelting and Refining Company and moved to consolidate it with Tottenville Copper. In 1914, as the US entered the First World War, Nassau Smelting and Refining was noted as one of four companies responsible for 90% of the country's copper exports.

It's said that war is good for business, and it certainly was for Lowenstein. The war brought a pressing need for copper, and the Tottenville plant was pressed into military service. This part of the company's history is, unfortunately, well-documented due to a scandal all too familiar to our present times: on March 27th, 1918, police officers seconded to Naval Intelligence raided the Tottenville copper plant and arrested sixteen laborers—Germans and Austrians, many of them crew members of German merchant ships who had become trapped in the United States by wartime turmoil. In finding productive employment, a way to support themselves, they had made the critical mistake of taking jobs that supported the war effort.

The intelligence officers spent most of the day at the plant in Tottenville, which employs about 500 workers. Practically all of them were questioned, but most of them were found to be either native Americans or naturalized citizens. The sixteen who were unable to show either that they had registered or had obtained zone permits were placed under arrest at the plant. (Q3)

Because of its role in supplying copper components of artillery shells, the Tottenville smelter was considered a munitions plant, and was thus off limits to any enemy aliens who did not possess a specific movement permit issued under police supervision. A few of the sixteen arrested had not registered as enemy aliens, but most had registered and had the wrong work permit. The newspapers do not suggest that these sixteen had committed any offense other than a lapse in paperwork, but they were nonetheless "turned over to Federal Authorities for internment." The authorities were also, reportedly, investigating an allegation that a manager at the plant had made "seditious remarks." These included criticism of the Liberty Bonds used to fund the war effort, and a suggestion that American forces in France would not prevail. Fortunately for the plant manager, he was able to produce paperwork proving his citizenship, and officers deemed the evidence of his "seditious" opinions to be insufficient for charges.

If the war brought good fortune to Lowenstein, peace took it away. The end of steady military contracts complicated the finances of Nassau Smelting and Refining, requiring a retooling of the plant towards other products in the difficult context of the post-war recession. A major fire at the plant, in 1923, racked up a huge repair bill and cut into production. There were personal problems, too: in the mid-1920s, Lowenstein divorced, starting a bitter multi-year legal battle over custody of his children—a question ultimately resolved in his favor, but not without the involvement of the appellate courts and an axe-wielding deputy sheriff. The financial condition of his company continued to decline as the country slid into the Great Depression. Lowenstein must have been looking for an exit.

By this time, Nassau Smelting and Refining was consolidated into a 45-acre property on Staten Island straddling Mill Creek, just south of State Route 440 and between Arthur Kill Road and Page Avenue. There were two primary operations: the "red metals" complex which processed copper, and the "white metals" complex for lead and tin. Both ran primarily on reclaimed scrap, refining it into ingots ready for reuse. Conveniently, these were two categories of metals in great demand to the Bell System: copper, for wiring, and lead and tin, extensively used to coat cables and splices and as key ingredients in solder. The post-war period brought not only general economic decline, but also an increase in metals prices, stressing AT&T's supply chain. Western Electric turned its mind towards consolidation. Given the Nassau plant's proximity to WE's headquarters in New York City and plants throughout the region, WE must have already done quite a bit of business with Lowenstein's company. In 1931, they bought it.

Likely because Nassau Smelting and Refining was already a well-established business, WE left it to operate as an independent subsidiary, alongside the Teletype Corporation (which was acquired at nearly the same time) and, later, the Sandia Corporation in Albuquerque—the three independent subsidiaries of Western Electric through the 1980s.

Western Electric promotional graphic

By the 1940s, Nassau Smelting and Refining was processing thousands of tons of scrap each year. Most of this was disused telephone equipment and cable that had been broken down at other WE plants and then delivered to Staten Island for smelting. WE reported that about 3/5 of the nonferrous metal content of this scrap was returned to WE as high quality metal stock for manufacturing use. In the mid-1950s, Nassau Smelting and Refining provided about 16% of the Bell System's copper supply and 20% of its lead. During the Second World War, when copper became exceptionally scarce, the Nassau works provided a critical in-house metal recycling capability that not only supplied the military as a contractor but also allowed AT&T a reliable source of metal for its wartime telephone projects. During several such periods of disruption in the metal market, Nassau Smelting and Refining provided most of AT&T's supply. For AT&T, vertically integrating metal smelting thus had two key advantages: cost savings from owning its own supplier, and a degree of protection from the whims of the market. "Nassau Smelting and Refining Company is a further extension of Western Electric's constant effort to do its Bell System job better and more economically" (Q2).

A 1946 newspaper article, announcing an open house at the plant with tours open to the public, gives a sense of the scale of the operation. Each day, an average of five railroad cars of scrap arrived at the plant. Stripping machines separated lead sheathing from telephone cables "like a child would peel a banana" before the remains were fed into the furnaces—enough oil to heat a house for a year kept each furnace at 2,000 degrees for one day. Lead was transferred from furnaces to kettles, 30 tons at a time, and poured into ingot molds. Much of the lead then went to the plant's on-site solder mill. Each of the presses there formed enough rosin-core solder to reach "from Tottenville to St. George [at the far end of Staten Island] and back," each day.

During the mid-century, WE expanded the Nassau company's remit beyond just nonferrous metals. Nassau Smelting and Refining became WE's general broker of scrap and secondary materials, brokering cinders from telephone company power plants as a soil amendment, and iridium recovered from telephone relay contact points as a precious metal. As a subsidiary, rather than a mere component of Western Electric, Nassau was somewhat more independent of AT&T than the rest of WE. The metal operation wasn't restricted to the telephone industry, and both purchased scrap on the open market and sold metals to any buyer. Nassau was one of two bidders, for example, on an enormous post-war Naval copper supply contract.

Metal refining is not a clean operation. In 1947, Nassau faced charges of "smoke annoyance" and "noxious conditions" that might have led to a criminal prosecution, were the case not forestalled the company's agreement to install a $350,000 bag house to filter furnace emissions. By this time, Nassau was called one of the nation's largest "above-ground mines." Metal recycling, while economical nearly from the beginning of metallurgy, received a surge of interest in a post-war nation that keenly remembered the shortages of the previous decade. Unlike our modern association between recycling and environmental protection, in the 1940s it was styled mostly as a new form of extractive industry: "The Nassau Smelting and Refining Company... mines the vast Bell network for valuable metals.... in 1948,.. Nassau reclaimed more copper than was produced in six of the nation's 14 major-producing states. And of 22 major lead-producing states, only four produced more than was reclaimed by Nassau."

The president of Nassau at the time, William Scheuch, described "American homes, factories, and cities" as the "mines upon which modern industry depends." He noted as well that, by that time, rubber, plastics, rope, and "scores of other materials" had fallen under Nassau's responsibility. "So thrifty are Nassau's experts that even floor sweepings are cooked by incandescent heat to reclaim the last drop of metal content" (Q4). Later that year, Nassau hosted a delegation of European metallurgists discussing techniques for recycling aluminum, then a major focus of Nassau's research department.


I put a lot of time into writing this, and I hope that you enjoy reading it. If you can spare a few dollars, consider supporting me on ko-fi. You'll receive an occasional extra, subscribers-only post, and defray the costs of providing artisanal, hand-built world wide web directly from Albuquerque, New Mexico.


In 1956, the Staten Island Advance ran a puff piece on Nassau's production (162 million pounds of scrap converted into 138 million pounds of saleable metal in 1955) immediately next to the headline "Smog Growing Problem." In 1958, the same newspaper carried an editorial by Nassau's new president, Arthur Fegel, under the headline "Nassau Battles Air Pollution." In some ways, the piece is a celebration of the company's 75th anniversary, but the headline betrays an underlying political struggle. "Just as good families are good neighbors,.. an industrial concern like Nassau devotes much attention to being a good citizen in its own neighborhood" (Q5). This was the preface to an announcement, at the end of the article, that Nassau was about to build a new furnace that would burn the insulation off of wire. The exhaust, he promised, would be "clean as a whistle."

In the early 1960s, Nassau underwent a wave of expansion, including new office and warehouse facilities on the north side of the property. The plant had become one of the major employers of Staten Island, with a payroll including "12 fathers and their 14 sons." The Staten Island Railroad operated a train station for the plant's employees, called Nassau, and expanded it in the early 1970s. The plant was one of their largest freight customers as well, with an industrial siding just off of the station. In 1971, the Deputy Commissioner of the New York Department of Air Resources paid a visit, or rather a "sniff," in part to confirm that Nassau had complied with an order to stop burning lead away in furnaces. "In terms of other smelters in New York City, it is fairly clean; but that doesn't mean it is in compliance with all the standards" (Q6).

The 1970s were, for the Bell System, the beginning of the end. An upstart subsidiary of the Southern Pacific Railroad waged an intense legal battle against AT&T's monopoly, won the right to compete on long-distance service, and renamed itself to Sprint before merging with principal AT&T competitor GTE. Corning demonstrated a new communications technology based on light trapped in glass fibers; by 1980 the manufacturing technique for these new cables had become refined enough that they presented a lower-cost option compared to AT&T's coaxial and microwave network (Bell Laboratory's alternate plan for the future of communications, long-distance microwave waveguide, was stillborn).

At the same time, the American environmental movement hit its stride. The Clean Air Act, the Clean Water Act, the establishment of the Environmental Protection Agency; each of these steps imposed new requirements on an aging metal plant that had long been considered a major polluter. Wastewater from metal separation processes, a slurry of heavy metals and petrochemicals and God only knows what else, was redirected from Mill Creek to the site's first water treatment plant in 1973. That plant separated the contaminants into a dried sludge, which for years was simply piled up under the approach ramp of the Page Avenue bridge. Later, this sludge was processed to extract precious metals, but that modest revenue didn't make up for the capital investment. It wasn't a good time to be an expensive part of Western Electric: facing declining revenues, labor unrest, and the weakening state of its sole benefactor AT&T, WE entered the late 1970s as a company in decline. WE started backing away from its integrated salvage operation: sometime in the late 1970s, the remaining metal recycling operations at Nassau were contracted to a company called C&D Recycling, which assumed management of part of the Nassau plant.

Copper smelting operations at Nassau ended in 1981, beginning a multi-year decommissioning and demolition process for the Red Metals complex. The main copper processing building was razed by 1985, but in the mean time, the entire Bell System had met a worse fate: divestiture. Antitrust lawsuits against AT&T, brewing throughout the 1970s, leading to a 1982 settlement agreement that required the dismantling of the Bell System over the following years. On the first day of 1984, the Bell Operating Companies became independent corporations, over 2/3rds of AT&T gone overnight. Western Electric went along with them: one of the key findings of the antitrust case was that AT&T had built and maintained a monopoly through vertical integration so extensive that it deprived their competition of equipment and supplies.

AT&T was able to partially mitigate the unwinding of their vertical integration: by agreeing to completely divest the operating companies, it won terms on which Western Electric could remain an AT&T company, under a new name. AT&T had agreed to end use of the trademarks most associated with their nationwide monopoly, including not just "Bell System" itself but the Western Electric name and logo. WE was reorganized, with some divisions meeting other fates, but most of the company became AT&T Technologies. As part of the settlement, AT&T had gotten relief on various restrictions imposed on them by previous antitrust cases, including a key prohibition on Western Electric marketing general-purpose computers. While there was some initial optimism that that small victory would allow the new AT&T Technologies to take on the likes of IBM, it didn't work out that way. A series of poor decisions, several fundamental missteps, and no doubt some plain bad luck had Western Electric, and its close partner Bell Laboratories, on the downhill.

Post-divestiture, Nassau Smelting and Refining took on a new identity: AT&T Nassau Metals. As part of the settlement agreement, Bell Operating Companies could no longer be required to purchase equipment and supplies from WE. The Bell telephone market was suddenly open to competitive manufacturers such as the Canadian company Northern Electric (later Nortel), itself a fragment of Western Electric that had broken away when a 1950s antitrust case led to a settlement agreement that WE would divest its foreign operations. On top of the Carterfone decision and divestiture making consumer telephones a competitive market, WE found its business seriously undermined. The 1980s saw closure of many of WE's largest plants, Nassau not excepted.

The White Metals complex at Nassau continued longer, shutting down in 1991 with demolition starting in 1996. The last manufacturing operations at AT&T Nassau Metals, by then limited only to electroplating, ended in 2001. Nassau had actually outlived its parent, WE, which was renamed to Lucent Technologies and made independent in 1996. By that time, AT&T Nassau Metals had been renamed to simply the Nassau Metals Corporation, and existed primarily to manage the closure and remediation of the Tottenville site.

All of the original, 1930s-era manufacturing buildings were found south of Mill Creek and had been demolished by 2000. The newer 1960s era buildings, an office building and a warehouse north of Mill Creek, were leased to a developer who found various commercial tenants. Most of the site remained abandoned in its post-demolition state for the next twenty years, though: a formidable environmental recovery was required before reuse.

The Nassau Metals site was investigated by the EPA for inclusion on the National Priorities List as a superfund site during the 1990s, but was ultimately not nominated. The main reason was simple: the Potentially Responsible Party was the Nassau Metals Corporation, a subsidiary of AT&T, a company that was still very much alive. With some cajoling by environmental authorities, AT&T agreed to avoid the federal CERCLA process by entering the New York Department of Environmental Conservation's Voluntary Cleanup Program (VCP). Many environmental authorities offer something like the VCP: one of the reasons that contaminated industrial sites, often called brownfields, tend to stay that way is the uncertainty and liability involved in environmental contamination. Real estate developers are understandably hesitant to commit to a property that may require an enormously expensive remediation in the future.

The VCP provides an alternative: when a company participates in the VCP, they develop a comprehensive plan for site remediation—at their expense—that is mutually agreed with and supervised by the Department of Environmental Conservation. In exchange for the land owner completing the approved remediation plan, the Department of Environmental Conservation makes a binding agreement not to impose further requirements in the future. In other words, it's a legal arrangement to settle on what level of cleanup is "good enough," so that future owners of the property are protected from expensive surprises.

For remediation purposes, the Nassau site was divided into three Operable Units. The largest, OU1, includes the primary industrial area where both the Red Metals and White Metals facilities had been located. From a 1991 Site Investigation Report to the 2011 Final Engineering Report, remediation contractors identified extensive contamination of the soil throughout the site, and downstream on Mill Creek, with lead and other heavy metals. It was found, for example, that a substantial portion of OU1 was built on artificial infill of a former wetland. The fill material, in line with standard practice in the 1930s, is best described as "assorted trash." Everything from construction debris to domestic garbage to old telephones had been piled up and compacted, and then factory buildings put on top of it all. Every time there was a storm, Mill Creek surged against its south bank and washed some of it away, downstream, into Arthur Kill. Particulate contamination from these materials could be identified out into the ocean.

During the 2000s, contractors dredged Mill Creek and parts of Arthur Kill, temporarily dammed Mill Creek to facilitate further excavation, stabilized the banks of Mill Creek by geotechnical methods, developed new wetland areas on other parts of Arthur Kill at a 3:1 ratio to the area permanently disturbed by the site, replanted vegetation, and cleaned out storm sewers that had been contaminated by runoff from Nassau.

These were all secondary efforts, though, in comparison to the largest remediation activity. The majority of the site, on both sides of Mill Creek, were covered with an engineered barrier of soil, stone, geosynthetic clay liner, and asphalt, intended to ensure that the contaminated soil will remain on site. It is simply too large of a volume to practically be removed, and even if it was, the degree of the contamination is so severe that it would be difficult to find a facility licensed to dispose of it. We tend to think of nuclear waste in the most severe terms, a contaminant that we can never be rid of, pretending that this is somehow an unusual outcome. The reality is much worse: permanent on-site entombment is one of the most common fates of industrial contamination, and there are tens of thousands of sites across the United States in which it is forbidden to dig.

Nassau is one of those. Nassau Metals no longer owns the land, it has all been sold to private developers. On the north side, where the office building and warehouse remain (now a restaurant and a complex of fitness facilities respectively), the parking lot and foundation slabs of two new fast food restaurants make up the permanent barrier. Whenever any paving or concrete work is required, a qualified environmental engineer must be on-site to supervise the work and ensure that the integrity of the barrier is maintained. On the south side of the property, which was unpaved, the clay barrier must remain in place under the new construction or a new barrier must be approved by the Department of Environmental Conservation. No digging can be done without arranging for disposal of the disturbed soil at a licensed facility. Covenants on the property require that these institutional controls be observed in perpetuity:

the owner of the property shall prohibit the Property from ever being used for purposes other than for Commercial or Industrial use without the express written waiver of such prohibition from the Department

Because of the particular standards to which the site was remediated, it is considered suitable only for limited occupation within the context of the institutional controls. Healthcare facilities, elder care homes, child care facilities, agriculture or gardening of any kind (even at hobby scale), and residential construction are all perpetually prohibited. The groundwater cannot be used without construction of a treatment facility approved by the state. The barrier system containing the contaminated soil must be inspected and recertified by engineers on a regular basis.

This, then, is the story of Western Electric—it calls for many things and many activities, blended together, to create the miracle of telephonic communications. It takes people and equipment, inventiveness and technical skill. It takes experience, and, above all, the fundamental desire to be of service to the public and the nation. (Q2)

For some 75 years, AT&T was the telephone company, and Western Electric made the telephones. It destroyed them, as well: every phone inspected, tested, and if found wanting, condemned to the furnaces of Nassau. Telephone wires stretched across the nation by 1915; and by 1984 that very same wire had likely been cut, stripped, melted, and refined on Staten Island; quite possibly, it had even been returned to Western Electric, just a few molecules in each of the 35,000 telephones rolling off the line of the Indianapolis Works. Likely also, a few molecules of the soil permanently interred under Premiere Pickleball of Staten Island, a few molecules on the banks of Mill Creek, a few molecules of the ocean.

Ingots at the Nassau plant

Metal is highly recyclable; it's sometimes estimated that the majority of the copper ever mined is still in use today. Much of that metal passed through Nassau at some point, there might be a bit of a Western Electric 500 in your smartphone today. There certainly is under the YoYo Chicken: a tombstone for the telephone era.

Error'd: Fi fa foe

First up this week is a little story about a fifafail. I do wonder if this was a failure of the television station, or whether there was something more to it than that.

Hercules wrote to alert us to these World Cup shenanigans, explaing "At least the flags were correct. And yes, this was live TV. The host got the country names correctly, and even called out that the written text was wrong"

5503f1d7141948f88f4650c17468ff46

"I'm very open in my job search but I did limit it to France. The search has been working well for months, but this morning I got a bevy of new interesting propositions. It seems France is much bigger than it was yesterday." Apparently WorkerNumber29200 is surprised by the expansionist nature of an imperialist coloniser. Plus ça change, Worker.

f05fd09185bd4f2ebeacc92de9609131

We have a couple of wtfs from Github. First Hans K. "would love to find a, so I could fix this GitHub Dependabot issue."

4d41995a2e7a423cbf8de6e960b876c4

And Peter S. figures that "GitHub has trouble doing basic math -- or they have an unpublished proof that 0=1"

e954781458d64f6da02bfad014efc854

Finally Michele has just encountered one of the most maddening phenomenon on Amazon recently. "Searching for a cheap USB-C fast charger. Got a list of expensive CDs of obscure artists." All of them AI-generated, like the 100000 Whys books?

cdf7c2baae7541a9ad8e90b9c4ecb0c4

[Advertisement] BuildMaster allows you to create a self-service release management platform that allows different teams to manage their applications. Explore how!

The Roadmap

When Gary was called in for a meeting with a few of his managers- because of course he had several- he thought it was going to be for an "attaboy", because things had been going really well for the past few months.

Gary had inherited a mess, and taken over a nightmare application. It was the kind of application that should be a simple CRUD-style data-driven app, but somehow despite only having 20ish entities it managed, someone had generated 500+ controllers for managing them. Most of those controllers were copy/pasted code with minor changes in the WHERE clause of a SQL query.

And that was just the code. The infrastructure was similarly a mess, with duplicate resources provisioned in their cloud host. There was no CI/CD, no unit tests to speak of, no deployment process that wasn't "manually copy these files and pray". And uptime? You've heard about "five nines", but this product was lucky to get even one nine. Especially because the manual deployment process meant a few hours of downtime.

And that was just the infrastructure. The backlog was similarly messy. There were lots of tasks- many thousands- but not a single one had a priority. Most of the tasks were something like, "Fix database timeouts", or "Bug 531" with no description to explain what they were. At best, some of the "new feature" tasks linked to a Google Doc that explained a software roadmap that had been last updated in 2020.

So with no guidance, Gary and the rest of his team got to work. Cloud costs were massive. Just cutting the duplicate resources would help, but with actual planning it wasn't hard to find even bigger wins. In total, Gary got the cloud costs down 60%- essentially saving the company a small multiple of his salary every year.

With that out of the way, getting a CI/CD pipeline running was next. Within a few weeks, manual deployments were gone. Everything was automated. Downtime nearly vanished. And now, with all the cost savings in cloud resources, for a fraction of what they were paying, it was easy to automate provisioning test environments for each new feature.

So Gary was very ready for some congratulations when he sat down with management. He was prepared to discuss all the wins he and the rest of the developers on the project had gotten over the past few months.

"I'm sure you know why we're sitting down," Manager the First said when they settled into the conference room.

"I'm sure," said Manager the Second.

"We have some concerns about your performance," Manager the Third said.

"My performance?" Gary asked.

"Yes," said Manager the First. "Let me pull up the backlog."

"And the roadmap," said Manager the Second.

"Yes, I'm getting that up too, thank you." The trio of managers struggled with pulling up the appropriate pages, and after about 15 minutes, gave up. Instead, they discussed their complaint without visual aids. "You haven't completed any of the tasks on the roadmap. Bug 673 has been open since you started on the team. None of the roadmap milestones have been touched. There's absolutely no progress."

"Okay, but that document was wildly out of date," Gary said. "Instead I put cycles into solving the actual problems we're having. I've saved the company a huge amount of money. I've gotten our development cycle time down to a fraction of what it was. And we have basically no downtime!"

"That's all very nice, I'm sure," said Manager the Third. "But none of that was on our roadmap."

"Well, maybe we should set up a meeting to go over the roadmap," Gary said. "Because a lot of the tasks on there don't make much sense right now-"

"I don't think that's a good use of time," Manager the Second said. "Large meetings are expensive. Just stick to the roadmap, please."

With that, the meeting ended. Gary went back to work…

… updating his resume.

[Advertisement] Keep the plebs out of prod. Restrict NuGet feed privileges with ProGet. Learn more.
❌