This online event marked the conclusion of the Traceable AI Answers for Public Data project under the AI Learning Labs initiative, showcasing AI integration with open data portals. You can now watch the full event recording and explore all presentation materials.
It’s October, and the 2026 Virtual DLF Forum is almost here! On the 14th and15th, we’ll come together online for two days of learning, collaboration, and shaping what’s next for the field. In the coming days, we’ll share final Forum updates, ways to prepare, and opportunities to connect across the DLF community. We can’t wait to see you online! Be sure to register by tonight if you haven’t yet.
I’ll be in Spokane, Washington, for the Joint Council of Librarians of Color (JCLC) conference as an exhibitor with Sharon M. Burney, Program Officer and member of the CLIR grants team, from October 7-11, and I’d love the chance to meet up for coffee with DLF’ers while I’m there. If you’re nearby and available, send me an email.
With excitement,
-Shaneé
This month’s news
Last Chance to Register: Registration for the 2026 Virtual DLF Forum closes today, October 1. Register now to join us online! Can’t join us live? The keynote will be streamed live and available afterward on YouTube, so you can tune in or catch up later.
New Website: CLIR affiliate The Shared Print Partnership (SPP) has a new website! SPP is a federation of monograph and serial shared print programs across the United States and Canada. Explore the new site for information about available programs and resources.
Office Closure: The CLIR offices will be closed on Monday, October 12, to observe Indigenous Peoples’ Day.
This month’s open DLF group meetings:
For the most up-to-date schedule of DLF group meetings and events (plus conferences and more), bookmark the DLF Community Calendar. Meeting dates are subject to change. Can’t find the meeting call-in information? Email us at info@diglib.org. Reminder: Team DLF working days are Monday through Thursday.
Born-Digital Access Working Group (BDAWG): Tuesday, 10/6, 2pm ET / 11am PT.
Digital Accessibility Working Group (DAWG): Tuesday, 10/6, 2pm ET / 11am PT.
AIG Cultural Assessment Working Group: Monday, 10/12, 1pm ET / 10am PT.
AIG User Experience Working Group: Friday, 10/16, 11am ET / 8am PT.
AIG Metadata Assessment Group: Friday,10/16, 2pm ET / 11am PT.
DAWG IT & Development: Monday, 10/26, 1pm ET / 10am PT.
Digitization Interest Group: Monday, 10/26, 2pm ET / 11am PT.
Committee for Equity & Inclusion: Monday 10/26, 3pm ET / 12pm PT.
Climate Justice Working Group: Tuesday, 10/27, 3pm ET / 12pm PT.
Open Source Capacity Resources Group: Wednesday, 10/28, 1pm ET / 10am PT.
DAWG Policy & Workflows: Friday, 10/30, 1pm ET / 10am PT.
Yesterday Kevin Ford and I gave a presentation
(slides)
at the LD4 conference about the use of BIBFRAME “bounded descriptions” at the
Library of Congress (LC) and in the Blue Core project.
BIBFRAME data is RDF that gets serialized in many different ways. One of
the challenges of building tooling around BIBFRAME is that you don’t
necessarily know if you have to follow
your nose in the data (resolve URLs) to get some of the data, or if
it will be present in the data you have in hand.
This may seem like a small thing, but when you are building BIBFRAME
editors, or doing a large bulk loading operations (as LC and the Blue
Core project are doing) it is helpful to have it on hand. The
bounded description is a convention for making certain
information available in the BIBFRAME data.
The initial idea for the talk was to describe some interoperability
problems that we
experienced when loading data that had been designed for XML processing
in a RDF processing centered workflow. But rather than focus on those
specific problems it shifted instead to talk about the role of bounded
description for interoperability more generally.
The bounded description aside, one thing that fell out of this work for
me was the importance of having a JSON schema for
BIBFRAME. You might wonder why on earth you would want a JSON Schema
for BIBFRAME RDF data – wouldn’t that be overly prescriptive, since RDF
can be represented using XML, JSON, Turtle, N3, etc? The answer is yes,
and that is completely the point.
Following along with the principles of Linked Open Usable Data (LOUD), it
is extremely useful for tool building purposes to have a clear idea of
how the JSON-LD is structured. Even within JSON-LD there are many ways
of representing the same exact RDF graph. Working with this variety of
serialization formats requires tools to use an RDF parser. For some in
the Linked Data community this isn’t a big deal.
But most people working with data on the Web are used to loading some
JSON with their programming language’s JSON support, and interacting
with it like a data structure: hash maps, lists, strings, integers, etc.
If they can’t reliably do that it is a major interoperability problem.
If BIBFRAME is really going to replace MARC I think it’s important to
let the ontology and RDF serializations exist, while providing a strict
and easy to use representation for the many people who aren’t engaging
at that level.
The bibframe-json site is brand new. We are storing JSON-LD in the Blue
Core data store, and making it available via an API. But we aren’t
actively using this schema (or any other) for validating the data on the
way into the data store…at least not yet. The schema was developed
rapidly with Claude Code, mostly as a provocation for the LD4 talk. It’s
my (probably unpopular) belief that if we can’t represent BIBFRAME data
using JSON Schema then BIBFRAME really isn’t going to see widespread
adoption, ever. I’m sure that there ways for this initial version to be
improved. If you have ideas or suggestions please consider contributing
to the work, or get in touch with me directly.
I saw some recent
news going around about the Australian Internet
Observatory. Bergis
Jules, my collabrator on the (now ended) [DocNow] project alerted me
to it, and I saw it getting shared around in the fediverse. Bergis
reminded me that we had been talking about the importance of the
“archive downloads” that the GDPR required web platforms to make
available to users, in order to abide by the GDPR. At the
time we had been mostly focused on Twitter, but others like the National
Library of New Zealand were starting to collect
Facebook archives (Wayback link since it appears they may have
stopped doing this?).
The AIO seem to have recognized this too, and built a platform to
connect researchers and data donors to contribute their archives which
they have called Data
Download Packages or DDPs.
While the announcement is recent, it appears that they have been doing
a lot of work analyzing the DDPs from a host of platforms
(AirBNB, Facebook, Google, Instagram, Netflix, Spotify, TikTok, Uber,
and X). They’ve developed a taxonomy
and schema for the data to help researchers use the data they need.
One thing that is not clear to me is to what degree these schema are
used to provide access to the underlying DDPs. But either way this work
will be very helpful not just for their platform but for the research
community in general.
VDMC and Q.E.D. Systems teams at the Q.E.D. Center for Training and Development.
How do we really know when someone has learned a skilled trade? Is completing a task enough, or should we also understand how the person performed it, where they struggled, what demanded their attention, and how their performance changed as they gained experience?
These questions came to life during a recent hands-on maritime workforce visit to Q.E.D. Systems, Inc., by members of the Virginia Digital Maritime Center (VDMC) Workforce Innovation team. For the day, we stepped away from the roles we normally hold as researchers, developers, and workforce professionals and became novice maritime trainees. Instead of only discussing workforce training and performance from the outside, we experienced parts of the learning process firsthand.
The experience gave us an opportunity to learn several maritime skilled trades firsthand while also observing how skills are introduced, practiced, coached, and ultimately assessed. Combining hands-on participation with observation made the visit especially valuable for understanding both training and performance.
From the Learning Lab to the Maritime Olympics
VDMC Team 1 and Team 2 relay teams during the Maritime Olympics Challenge.
The day began in the Learning Lab, where we rotated through hands-on activities involving marine electrical work, rigging, welding, and fall protection. We quickly realized that learning a skilled trade is very different from simply watching someone perform it. We listened carefully to instructions, observed demonstrations, handled unfamiliar tools, attempted new tasks, made mistakes, received feedback, adjusted techniques, and tried again.
One important takeaway was realizing how much is involved in performing even a seemingly simple task well. We had to pay close attention to instructions, coordinate our movements, remember each step, make judgments as we worked, and continuously check whether the task was being performed correctly.
Later in the day, the Learning Lab became an Assessment Lab through Q.E.D.'s Maritime Olympics. The skills introduced earlier were no longer just something we had learned or practiced; we now had to put them to the test through a series of performance challenges and demonstrate what we could do.
Experiencing the Challenge as Learners
VDMC team member Lawrence Obiuwevwi is practicing welding during the hands-on training
Another activity that stood out to us was welding. It is easy to observe an experienced welder and focus only on the finished result. As we took on the role of learners, we began noticing everything that happens before that final result: where attention is directed, how the hands are positioned, whether the correct movement is maintained, how quickly an error is recognized, and how performance changes after feedback.
That experience raised an important question for us: How do we move someone from being introduced to a skill to being able to perform that skill reliably in a real working environment? And perhaps more importantly, how do we know when that transition has actually happened?
Looking Beyond Task Completion
Traditional workforce assessment can focus on whether the trainee completed the task, how long it took, how many errors occurred, and whether the final product met the required standard. Those measures are important, but participating in the activities reinforced that they tell only part of the story.
Two people may successfully complete the same task while going through very different processes. One person may move confidently and efficiently, while another may hesitate, repeatedly correct mistakes, require more guidance, or experience much greater cognitive demand. If we look only at the final product, much of that difference can remain invisible.
Within VDMC's Maritime Human Performance Science effort, led by Dr. Jessica M. Johnson at the Neuroergonomics and Simulation lab, we are interested in how multimodal technologies can help explain what happens during skilled performance, not simply what the final outcome looks like.
As we moved through the activities at Q.E.D., we began thinking about the kinds of questions that could be investigated using eye tracking,physiological sensing, neurophysiological measurement, behavioral observations, and performance measures. Where does a trainee look during a critical step? When does a task begin demanding more attention? Where do hesitation and uncertainty appear? How does a learner respond after an instructor provides feedback? And what begins to change as the learner becomes more proficient?
Those are the kinds of questions that move us beyond simply asking, “Did the trainee complete the task?” toward understanding how skill is actually acquired. This is where the field-trip experience connected directly with the research being pursued at VDMC.
Toward a Process-Level Understanding of Skilled Work
From a research perspective, skilled work is especially interesting because many different kinds of information are generated at the same time. Eye tracking can tell us something about visual attention. Physiological sensing can help us examine changes associated with workload and arousal. Neurophysiological measures can add another perspective on cognitive processes, while task-performance data provide the context needed to understand what was happening at a particular moment.
The challenge is not simply collecting all of these data. The important question is how we synchronize, organize, analyze, and model them in a way that produces information instructors and trainees can actually use.
The long-term goal is not to add more sensors simply because the technology exists, but to use these measurements to better understand how trainees learn, where difficulties emerge, and how proficiency develops over time.
One example is our ongoing pipefitting research in the lab, where we use a holographic training environment to study skilled-task performance. By combining task-performance measures with multimodal sensing, including eye tracking, physiological signals, and neurophysiological measures, we can examine how trainees respond to instructions, move through different task steps, encounter and correct errors, and adapt with practice. These insights can help inform better simulations, training approaches, and instructor feedback.
Sample trainee interaction with the Proto holographic pipefitting training environment.
Learning Maritime Skills Firsthand
The rope-knotting activity was another good example of why being physically present matters. Watching an instructor demonstrate a knot and actually trying to reproduce it are two completely different experiences. Reproducing the knot requires translating what was just observed into coordinated hand movements, while instructor feedback becomes part of the learning process. Experiences like this reminded us that it is difficult to fully understand skilled work from a laboratory alone. By taking the position of novice trainees—receiving instructions, using unfamiliar tools, making mistakes, correcting them, and later being assessed, we gained a different perspective on the complexity behind skilled performance.
VDMC team learning maritime rope knotting at the Q.E.D. Center for Training and Development.
Connecting Research, Training, and Workforce Readiness
One of the biggest takeaways from the visit was how closely workforce practice and workforce research can inform one another. We begin by identifying real workforce challenges, translating them into relevant research questions, collecting evidence, and using what we learn to improve training, assessment, and ultimately workforce readiness. For us at the Virginia Digital Maritime Center (VDMC), this is especially important because modern training environments can generate large and diverse streams of human-performance data. Our research increasingly focuses on understanding the process of skill acquisition, not only whether a trainee succeeds or fails at a task, but also how cognitive, physiological, behavioral, and performance patterns change as that trainee progresses toward proficiency.
This leads to an important research question: How can we move beyond measuring task completion and better understand the processes through which skilled competency develops? Answering that question requires collaboration across computer science, human-performance research, training, simulation, workforce development, and industry, with skilled workers and trainees at the center of the research.
We are grateful to Q.E.D. Systems, Inc. for opening its doors to us and allowing us to experience maritime training from the perspective of learners. We also appreciate the support of the Submarine Industrial Base Program Office and the broader maritime community working to strengthen workforce development and innovation.
We arrived at Q.E.D. as a workforce innovation team interested in training and human performance, but for part of the day we also became trainees. That shift in perspective was valuable because some research questions become much clearer when we step outside the laboratory, put on the safety equipment, pick up the tools, make mistakes, receive feedback, and experience the learning process firsthand.
Special Thanks and Acknowledgments
I would like to express sincere gratitude to Dr. Jessica M. Johnson, Director of the NeuroErgonomics & Simulation Lab and a leader of VDMC's Workforce Innovation efforts, for continued mentorship, guidance, and opportunities to participate in meaningful research at the intersection of human performance, workforce development, and innovation. We are equally grateful to Q.E.D. Systems, Inc. for sharing the expertise of its workforce and instructors and giving us the opportunity to learn by doing.
maritime human-performance research. We are grateful as well to the Virginia Digital Maritime Center, the Web Science and Digital Libraries Research Group, NIRDS Lab, and the broader Old Dominion University research community for fostering the interdisciplinary environment that makes this work possible.
Lawrence Obiuwevwi Graduate Research Assistant | Virginia Digital Maritime Center Department of Computer Science | Old Dominion University Norfolk, VA 23529 Email: lobiu001@odu.edu
Like most software developers, I devote some of my attention these days to what to make of the whole LLM thing. But because I’m a tech lead, I think about it a lot in terms of how it will impact my team. I, for example, value code review. I think it’s helpful for surfacing defects but particularly helpful for building shared domain knowledge and software practices across a team. But what happens to code review when a single dev can generate so many thousands of lines of code in a day? Current practices absolutely collapse under their own weight.
I’ve seen a ton of social media & blog posts about individual users’ techniques for navigating LLM coding, right? But they all read as “this is how I choose a model” “this is my fave harness these days” “here’s my top prompt hacks”, etc. — all about how single inviduals churn out code; fundamentally solipsistic. So I went looking1 for discussions on whether & how LLM coding can be metabolized by teams and didn’t find any.2 I guess that leaves me to write this post, for which, at the moment, I have more questions than answers. And insofar as I even have answers, someone is wrong on the internet and perhaps it is me, and thus I can harvest corrections!
And thus, some areas I’m wrestling with and possible solutions I’m mulling over. Stick around long enough and I’ll make a threat about legacy code.
Code review
A reviewer can’t usefully handle more than 200-400 lines of code in a review. One of my devs, experimenting, prototyped a 2,000-line feature in half an hour.
Do we, uh, have 5-10 code reviews? Who is doing those code reviews? Does a team need to be spending the actual majority of their time doing code review now? If so, does everybody hate their life?
This does not seem feasible.
Can you solve, or at least mitigate, the LLM-caused problem with different LLMs? I’ve been experimenting with a Claude skill to reorganize commit histories for big code dumps like this into reviewable units and it’s pretty okay, actually — but turning one 2,000 line PR into 10 200-line PRs merely converts the mostly-impossible to the possible without actually addressing the fact that everyone is now swamped.
I guess there’s automated code review tools. I’ve experimented only a little with them, mostly because when I tried, it sucked so much I immediately decided to spend my time on something else. But maybe there’s better stuff in the space I haven’t found.
I do think it immediately becomes important to beef up CI with all manner of automated (deterministic!) checks. I mean. Hopefully there was already CI checking stuff. But maybe you didn’t have a code complexity checker, and maybe you really really need one now. Test coverage checkers get super mandatory (on which more later). We can avoid even taking up reviewer time before a bunch of automated gates get passed. But it still seems like too much reviewer time. So….
Tests
Maybe code review just turns into reading the tests?? Particularly as the underlying code is disposable. Maybe it’s a waste of time to read said code (in detail or at all), as long as the tests pass.
But. Buuuuuuut.
Far and away the tests it’s going to be easiest to check are the integration tests, because we wrote the code at the level of integration testing: prompts which describe the overall behavior of a feature. We’re cognitively equipped to verify that the integration tests match the overall desired behavior in a way we’re not cognitively equipped to make sure the unit tests make sense.
But.
Integration tests kinda suck? I mean, they’re important, they’re the right level for testing certain things, sometimes you really do have to check that everything works together. But they’re slow and they’re brittle and they’re often smoke tests, rather than solid assertions that the underlying logic is consistently performed correctly. I have an important system in my life that has some integration tests and no unit tests3 and an enormously large fraction of my, and my team’s, time is spent dealing with this system. Building a world that rests on integration tests really sounds to me like building a world which will, sometime within the next 18 months, be a cannon pointed directly at my face.=
Also. Also?
LLMs suck at writing tests. Seriously, y’all. Perhaps you have One Weird Trick for avoiding this4 but in my experience they are simultaneously too verbose (in the way that LLMs are) and not verbose enough (missing important cases). Tests are where the stochastic-parrot-ness of it all show up the most, because to write well-chosen tests you really need to understand the nature and also the purpose of the system.
I could live with only reading the tests in code review, but they’d have to be good tests that really covered the bases. And it might be hard for me to tell what the bases were if I’d only grappled with the prompt and not the code. And if I left them up to the LLM, I wouldn’t trust them to be adequate.
Also, does anyone actually like reading the tests? Like isn’t that the part of code review people are most tempted to speedrun?
Observability
Cool cool, so maybe you don’t have unit tests and you have no idea how your code works.
Seems that it just got extra important to have pervasive and actionable monitoring. Not like that wasn’t already important. But are you going to know if your code, say, merrily opens up every available Postgres connection and never closes them? I mean. Not until you deploy it. Of course you could have missed that sort of thing in artisanal code too, but it’s a lot more likely you or your reviewer would have noticed.
Architecture and code quality
It’s good for a software team to understand 1. the architecture and key components of the systems it maintains and 2. the domain models within which it operates. Code review is a big part of how this gets spread! But that works less well and maybe not at all in an LLM world, per the above.
OK so. What tools are available here?
I think just talking about it gets even more important. I mean. It was always great to hang out and jam with coworkers about design ideas, data modeling, all that good stuff: what architectural decisions we’re going to make and why. But it gets more important if we no longer really have code review and we’re less often (or never) super deep in the code. Like maybe everyone’s schedule needs to intentionally have more meetings. I hear a lot of people groaning as they read. Sorry. Hold that thought because I’m about to make it worse: it might also get a lot more important to whiteboard stuff. And I’ve never actually found a tool for doing that which is anywhere near as good as a whiteboard…which means being in the same office together. (Groaning intensifies! I told you so.)
I think also it gets more important to collaboratively develop skills (like actual Claude skills, Copilot instructions, etc.). A good team will have more-or-less shared notions of things like…
what high-quality code looks like in their language(s) and framework(s) of choice
key ideas in the domain
architecture within and across codebases
important reusable components
favorite tools, dependencies, and/or patterns
….but they may or may not have documented these ideas. If LLMs are writing a lot of the code, I think it becomes much more important for the team to discuss and maintain (like, under actual version control) a set of skills which formalize all this. And then to point new devs at these skills during onboarding and have a clear expectation that everyone is using them. I think these git histories, conversations, et cetera get to be a really important mechanism not only for developing shared understandings but also for working them into the code.
Cognitive debt
Congrats, you’ve stuck around until I say that in an LLM-codegen-heavy world, all code is legacy code.
I can’t use the definition of “code without tests” here, because I just stipulated that we’re writing those, albeit in a super fraught way. I can’t even say “code we’re scared to change”, because LLMs have no such fear, and will churn out all manner of change! So I’m thinking in the sense of “code we do not understand”.
And yeah, I think code we don’t understand remains a problem in a world where we can write code super fast. Because code is not merely a set of tickets to be implemented, but also a model of a domain and a system for handling processes and all sorts of other things that actually require some sort of overarching Fred-Brooksian architectural understanding, something to provide conceptual unity and clarity across a broad scope.5
Because you can only go so far with implementing tickets or even features, right? Eventually you run into problems like “the fundamental architecture of this system collapses under our current load, and oh noes we want 10x more” or “I know I need a thing that is generally shaped like X but I don’t even know what things are shaped that way; I have to distill my understanding of the domain into something describable”6. And when you have problems like that, you actually need to sit down and think about things. What does my domain look like — what are its key nouns and verbs? What are the major components of this system, and how do they fit together, and are there seams where we can pull them apart and put them back together in a new way? If there aren’t seams, where and how do we make them?
I’ve had some luck telling LLMs to put particular kinds of seams in particular kinds of places, but I had to figure out what I wanted and how to phrase it first.
As with some of the paragraphs above, I think there’s some room for “use LLM tools to fix the problems created by LLM codegen” — huge shout-out to diffity-tour for helping me understand entire codebases in a radically accelerated way. But…
Meat
We’ve got these tools made of extremely eager linear algebra spamming the world with freshly laundered Stack Overflow, but our brains are still made of meat.7 And we made it faster to write code, but we didn’t make it faster to socialize ideas, or to build good mental models of new domains, or to hash out a design with a coworker, or quite frankly to manage the executive-function avalanche of identifying and allocating the number of tasks the robots can now do.8
I don’t actually think the tools we need for managing meat look like “how do be MORE productive! do the work of six devs by yourself! This is winning!”
I do think the tools we need for managing meat change when people’s jobs change, though. When we are suddenly needing to use different skills and tools. When we’re doing so in an atmosphere of hyper-precarity, both from geopolitics and economics writ large9 and from the rather obvious incentive to fire the other five devs who used to be doing that six-devs job, except maybe you’re one of the five, not the one.
Claude gets faster but meat does not.
By which I mean, I asked on bluesky, lol. This is not comprehensive research. ︎
I guess I could’ve also asked Mastodon and I’m sure I would have gotten more responses, but I’m also sure too many of them would have had me asking “why do I hate myself enough to have posted this on Mastodon”. ︎
There are good and interesting reasons for this, by which I mean “it seemed like a good idea at the time” and “may you live in interesting times”. But look, we know these things happen sometimes, much as we wish they would not. We are people of action; lies do not become us. ︎
The Brooks architect comes up in a context that’s entirely too into waterfall software dev, but it’s not like they had an alternative when he wrote it. And despite how much I dislike waterfall, I think he was right about the importance of conceptual clarity and unity in a codebase. Plus which if we’re writing a bunch of LLM code we are sort of secretly living in a waterfall world again, now aren’t we? Because we have just defined all the requirements up front and then sent our minions off to execute. So if we don’t know how to do that with the insight and elegance of Brooks’ architect, we are probably bad at things and should remedy our ways. ︎
I did this last year in a way that involved just kinda thinking about stuff for a month and turning that into a microservice that turned out to be pretty transformative for how the product works and what we can even do with it. The stuff I was thinking about included “graph theory” and “twelfth century Byzantine manuscripts”. Claude could never. ︎
It’s not just me, right? I’m not the only one who looks at LLMs and sees “oh, cool. So we have replaced a world of understanding and crafting things with a world of managing tasks, a world which has fundamentally higher and escalating executive function demands, but none of us are out here suddenly getting better at executive function and also software development is notable for specifically attracting people with executivefunctionchallenges and also in the US at least we’re run ragged from living under mad-king fascism, and so maybe a machine for increasing executive function demands is not actually our best move”? Anyone? Anyone? Bueller? ︎
Anyone who has interacted with me the last 3.5 months is likely to have heard that job searching is hard right now. Some describe it as a nightmare, saying that it’s harder than ever. There are so many stories of people who have been out of work for months and years. This post covers my … Continue reading "Job searching is hard right now"
Status quo costs are the direct costs, risks, and missed opportunities associated with allowing existing arrangements to persist. Several recent OCLC Research studies spanning a range of library-related topics have highlighted status quo costs as a key aspect of library decision-making. The takeaway from these studies is that status quo costs shouldn’t be neglected; recognizing them helps overcome inertia and makes the true cost of inaction visible in decision-making.
In this post, I’ll introduce a simple framework for thinking about status quo costs in library decision-making. Accounting for status quo costs helps libraries make more informed choices by bringing the often-hidden costs of continuity into view. By explicitly addressing the costs of maintaining current arrangements, library leaders can weigh the pros and cons of change versus continuity more realistically.
Status quo is a state of choice
The status quo is rarely an inviolable state of nature, or “the way things are.” Instead, the status quo should be understood as the accumulated result of past decisions based on conditions prevailing at the time: priorities, resourcing, technological capacities, and so on. Choices that were made under those conditions may have been optimal then, but are they still fit for purpose?
In the OCLC Research report Library Collaboration as a Strategic Choice, we frame the enduring influence of past choices using the concept of path dependency: the tendency for the status quo to continue because change is difficult to initiate. Path dependency is especially powerful when the choices involved in the creation of the current status quo gradually disappear from view. Over time, decisions lead to new routines, new routines become established practices, and established practices become unquestioned assumptions about how things should be. This makes the status quo seem inevitable rather than the result of earlier choices.
An anecdote from one of our earlier reports, Sustaining Art Research Collections: Case Studies in Collaboration, provides a compelling illustration of how path dependency, and the persistence of the status quo, can set in. A collaborative partnership between an art museum library and a nearby university library had, over the course of more than a decade, become routinized, and regular contact between stakeholders at the partnering institutions diminished. Routine, however, led to a lack of active management of the partnership, including regular evaluations of its progress or opportunities for adjustment or expansion.
This suggests an important lesson: decisions rarely are once and for all but instead should be revisited regularly as conditions and priorities evolve. This insight inspired one of the report’s recommendations: “One aspect of change management is having structures in place to regularly review the status of the collaboration and consider opportunities for expansion, course corrections, and other adjustments to the status quo. . . . Collaboration is an ongoing choice, not a one-time decision. Partners should continually and intentionally choose whether and how to collaborate.”
Whether the focus is library collaborations or something else, it is important to recognize that continuing the status quo is a choice that needs to be periodically reviewed and evaluated. This is a key first step toward loosening the grip of path dependency.
Surface the costs of the status quo
Once the status quo is visible, the next step is to make the costs of preserving it explicit.
All choices have costs. When the choice is change, the costs of adopting a new strategy are called switching costs. They may include necessary investments, steep learning curves, workflow disruption, and of course, the possibility of failure. These costs are very real, and library decision-makers rightfully take them seriously. In uncertain environments, switching costs often seem very apparent, even stark. In some cases, these costs are perceived as prohibitively high, producing a state of “lock-in”: the persistence of existing arrangements even when better alternatives are available.
Daunting as the costs of change may seem, the costs of preserving the status quo may be even higher. Maintaining the status quo can mean accepting ongoing inefficiencies, missed opportunities, and exposure to future threats. The challenge is to find a pathway through these competing considerations, a task which is often more art than science, shaped both by the magnitude of the costs involved and the decision-maker’s risk tolerance.
Accounting for status quo costs is vital for all types of library decision-making, including how to position an academic library within its parent institution. OCLC Research recently published the report The Library Beyond the Library, in which we make the case that academic libraries should engage more deeply with the broader institutional environment in support of research and learning. Increasingly, a library’s ability to fulfill its mission, retain influence, and demonstrate impact and value depends on that engagement.
For many libraries, this kind of engagement requires significant change in operational approaches. During an OCLC Research Library Partnership roundtable discussion on cross-campus partnerships in research support, many participants questioned whether a partnership-forward approach – an important pillar of the Library Beyond the Library framework – could work on their campus. They cited difficulties like the entrenched decentralization of the university environment and the pressure for individual units to demonstrate a distinct value proposition to campus leadership. These are genuine barriers to change, and overcoming them can involve substantial switching costs.
But the discussion also surfaced potential costs of continuing the status quo of providing research support services in isolation: diminished campus visibility; less impact; and ultimately, fewer resources. One participant provided a concrete example: on their campus, training workshops conducted in partnership between the library and another campus unit attracted far more attendees than workshops offered by the library alone.
By highlighting potential consequences of preserving the status quo, library decision-makers can see that maintaining continuity, while avoiding switching costs, nevertheless involves incurring other costs. And those costs can sometimes be higher than the costs of change.
Make the status quo a choice
Finally, make preserving the status quo an explicit choice. Place the status quo and its costs alongside alternatives (and their costs) to frame an intentional choice between continuity and change.
Once its costs are surfaced, the status quo is no longer a neutral state, but one option among others, with its own cost/benefit profile. This places the status quo on equal footing with alternative pathways—both need to be justified on the basis of their potential benefits and costs. It also counters the perceived inevitability of the status quo in the face of what might appear to be considerable switching costs for alternative options. The switching costs may indeed be high, but what matters is the comparison to the costs of maintaining the status quo. Evaluating these relative costs may reveal that continuity is the appropriate choice, but in this case, the status quo is intentionally chosen as the best option.
Making the status quo a choice could take the form of a periodic review, a set of established criteria triggering a change/don’t change decision, scenario comparisons, and so on. Consider this example from research data management. Once a data set is deposited into a repository, the tendency is often to let it remain indefinitely. But storage space is not infinite, stewardship consumes resources, and so the status quo of retention may not be optimal. The Illinois Data Bank counters this inertia by setting up an explicit choice about whether to continue retention. The Data Bank established a policy stipulating that retention of a data set is reviewed five years after deposit according to a defined set of criteria. Continuing the status quo—i.e., continued retention—therefore becomes an intentional choice within a clearly structured decision framework.
Another illustration is found in OCLC Research’s recent case study of the Portage Network, a collaboration among Canadian libraries to develop shared infrastructure and expertise for research data management. This included the creation of the Borealis Canadian Dataverse Repository, a shared repository that would supplant a number of existing local Dataverse repositories. In these cases, institutions made an intentional decision to move away from the status quo in favor of an alternative method of sourcing RDM capacity. The case study concludes that the status quo costs of maintaining a local Dataverse instance exceeded the switching costs of moving to shared infrastructure. What were the status quo costs? The cost of allocating staff to maintain local infrastructure was certainly a factor, but other costs were the levels of quality, efficiency, and speed foregone by retaining local repository implementations.
Making the status quo a choice is not a way to favor change over continuity. It is instead a way to avoid the pitfalls of path dependency, or a “this is how we’ve always done it” organizational ethos that may obscure better alternatives.
Counting the costs of continuity and change
The landscape of resourcing, staffing, technology, stakeholder needs, and much more is constantly evolving around libraries. In a dynamic and uncertain environment, library decision-makers need to recognize and evaluate both the costs of continuity and the costs of change.
What was once a choice can become entrenched as the status quo through time and familiarity. Making this visible, surfacing its costs, and deliberately reconsidering it alongside other options restores the status quo to what it should be: a decision.
In challenging environments, with limited or no internet connectivity or
in the face of repression, Tella makes it easier and safer to collect,
protect and hide sensitive data.
AI_CONTEXT is a lightweight, vendor-neutral pattern for preserving the
engineering context that normally gets trapped inside one AI chat, one
coding agent, or one context window.
The idea is simple:
Put the important project memory in the repository, then let every AI
agent read and maintain the same source of truth.
Instead of depending on a particular model’s conversation history,
AI_CONTEXT stores concise, durable project knowledge alongside the code:
today, I want to share a concept that’s helped me tremendously in this
shift. I call it micro-practices — 3 minutes, 7 minutes, 12 minutes — as
tiny, almost negligible containers of time that allows me a way into my
creative practices while sideskirting the resistance coming from my
brain that might say… no no no, but we’re too busy/tired/distracted/in
the wrong headspace for this right now…
To be a technology worker in a library, especially a publicly-funded
one, is to be subject to somewhat different pressures than those who are
employed by traditional technology companies. The financial pressure is
not worth lingering on – we know that we are paid less than those who
work for major corporations, but accept it as an inevitable result of
choosing to sell our labor to an organization that is not (at least
officially) oriented toward profit above all else. The mission of a
library is instead – in theory – the creation, preservation, and
transmission of knowledge, and those of us who appreciate that mission
may find ourselves drawn to this work
What does it take to help shape one of the most consequential
revolutions in tech history? Join Boris Cherny, creator and head of
Claude Code at Anthropic, for a wide-ranging conversation about the
past, present, and future of artificial intelligence. From the
challenges of building models like Claude to the possibilities ahead,
Cherny will offer an insider’s perspective on a technology rapidly
transforming our world.
Every request’s energy is measured directly from GPU hardware, not
estimated. Every carbon number uses the live carbon intensity of the
grid running your workload. No annual averages. No projections. We
report the physical carbon intensity of the grid our servers actually
run on, not the carbon we could claim by buying renewable energy
certificates elsewhere. The atmosphere doesn’t distinguish between a
certificate and a smokestack. Neither do we.
This led to the natural question: can gzip do language modeling?1 No
neural network, no learned parameters, nothing. Just the compressor that
ships with your operating system. You prime it with a corpus, give it a
normal text prompt, and it continues that prompt by searching for the
byte sequences that compress best.
Stalker (‹See TfM›Russian: Сталкер, IPA: [ˈstaɫkʲɪr]) is a 1979 Soviet
science fantasy film directed by Andrei Tarkovsky with a screenplay
written by Arkady and Boris Strugatsky, loosely based on their 1972
novel Roadside Picnic. The film tells the story of an expedition led by
a figure known as the “Stalker” (Alexander Kaidanovsky), who guides his
two clients—a melancholic writer (Anatoly Solonitsyn) and a professor
(Nikolai Grinko)—through a hazardous wasteland to a mysterious
restricted site known simply as the “Zone”, where supposedly exists a
room which grants a person’s innermost desires. The film combines
elements of science fiction and fantasy with dramatic, philosophical,
and psychological themes.
SearXNG is a free internet metasearch engine which aggregates results
from various search services and databases. Users are neither tracked
nor profiled.
For one, AI is not trained on what it means for code to be maintainable.
For instance, any reinforcement learning done needs a reward signal that
can be measured immediately, not in months or years. The AI learns rules
from rulebooks meant for beginners. The AI notices patterns from code in
the wild and let’s be honest, most code in the wild is pretty bad. There
is no fitness function you can define for maintainable code, at least
not one that we can discern, otherwise it would’ve been baked into our
linters.
Spain’s Second Section of the Intellectual Property Commission, a part
of the country’s Ministry of Culture, has ordered the blocking of
several domains of the Archive.today service, a web archiving service.
It means that most Spanish internet users who try to access the site are
redirected to a government page telling them that they are trying to
access an “illegal” website.
The last document from the Snowden archive was published on 29 May 2019.
The Guardian stopped publishing documents in February 2014, Der Spiegel
in January 2015, and The New York Times and ProPublica in August 2015.
After that, only The Intercept was still publishing documents, with only
a few exceptions, until it closed its archive in March 2019. Eleven
weeks later, on 29 May 2019, it released what would become the final
batch of documents from the archive. Since then, no news outlet,
journalist, or institution anywhere has published a single document from
the Snowden archive.
A Core Preservation Process is a specific action that every Trustworthy
Digital Archive should undertake adequately - either directly or through
its associated parties or services, in order to fulfill its digital
preservation missions as evidenced in its preservation policy.
Why, and how, are any of us capable of self-reflective thought? Of
experiencing emotion, ratiocination, preference, cold, whim, melancholy,
joy? Anything at all?
Userscripts run in a users web browser and make on-the-fly local changes
to specific web pages. In MusicBrainz they are generally used to change
the display of pages, facilitating editing.
Earlier this month, I added a post here about Five Thank Yous...this year's summer vacation project.
This weekend I was working on last summer's vacation project and realized I hadn't posted about it here.
That project is This Was News, a website and social media bot that encourages people to reflect on the news that was.
It seems (to me, at least) that the news cycle is winding faster and faster, which makes it easy to forget what the most important story of the day when it happened.
One reasonable reaction to this is to tune out the news entirely.
Another reaction would be to get periodic reminders of those stories, then be able to think about their impact after time has passed.
Although this site first debuted in 2025, its origins go back to February 2024 when a collision at DC's National Airport between a military helicopter and a commercial jet was the biggest news.
This was, of course, a big deal, but I wondered how soon it would fall off the front page and eventually be forgotten.
(Do you remember the around-the-clock reporting on that tragedy?)
It wasn't long after that before I wrote the "gatherer" stage of the project: an automated process to save lead article titles and URLs from the home pages of the Associated Press wire service, NPR, and the New York Times.
It ran for several months while I occasionally worked on other parts of the project.
(Eventually, I wrote a "back-fill" script that pulled the headlines from the Washington Examiner and Center Square websites to round out the political biases.)
The "Gather" script runs twice daily: 6:30am and 6:30pm Eastern U.S. Time.
That felt like the right time to capture the important news of the day: early morning East Coast time to get the important news from the previous day, then half a day later to see what happened in the first part of the current day.
The Gather script marks the "top" story from each site, then saves the next two to four stories in the order in which they appear in the HTML code of the page.
I'm using the order of appearance as a proxy for how important each news organization thinks the stories are.
The Editorial Policy page describes how I pick the news organizations and which stories are used from each site.
The way I envisioned using this was to get periodic posts in my social media feed, then replying to the post saying "remind me in 30 days".
Then a month later the bot would send me a personal message reminding me about the event.
As it turns out, that "remind me" function isn't as useful as I thought (and people haven't been using it) but it is still there.
A good number of people are subscribed to the Bluesky and Mastodon accounts, a few have subscribed to the daily email, and an unknown number use the RSS feeds.
(All of these methods are described on the "Get Reminded" page.)
The site's BlueSky and Mastodon bots will also record comments, likes, and reposts on each day's page to create a running commentary on the news.
Last summer's vacation project was a push to get all of the pieces functioning as a system.
I tried a "quiet launch" for ThisWas.News...I was curious to see how far it would go on its own without me pushing it with my existing (meager) social media presence.
For the first six months or so, it got a couple dozen people following the site's bots and reposting its content, which I thought was an interesting organic spread.
Since then I've posted about it a few times with my Bluesky and Mastodon accounts, and reshared some of its content, so the project is quite firmly tied to me now.
Fall 2026 Update
The site has been gathering news for about two and a half years and has been publicly available for about a year.
Over that time, I found that one of the news sources wasn't very useful, so I dropped Center Square and added the Wall Street Journal's news page.
(I had initially rejected the WSJ as mainly a finance news source, but reconsidered this year.)
I also added The Guardian from the U.K. and National Post from Canada as news sources with a bit of an outside-U.S. perspective.
The difficulty has come with the anti-webscraper features that news sites are putting in place now.
I'm not going any deeper than the front page of each news site to get the title and link to the article, but even that was too much for the New York Times and Associated Press.
I reluctantly switched to parsing the New York Times' RSS feed.
This isn't great because the RSS feed doesn't accurately reflect the order of articles in the page's HTML, so it feels like a derivative source instead of what the NYT editors think is really important.
There isn't an alternative for the Associated Press at the moment — they don't have a public RSS feed, and it seems like they are pushing people toward their developer API.
Needless to say, I'm not willing to pay for a developer API key for a solo, spare-time project like this, so I might drop the Associated Press.
(If anyone has another way to get the lead stories from the Associated Press, please let me know!)
A few months ago, I also added "tags" to the data model for each story, and those get output as linkable tags in the social media feeds.
That has helped with discovery of the posts in the social media feeds quite a bit.
The core idea still holds: the news moves fast, and it's worth having something that slows it back down.
Two and a half years in, the earliest stories the Gather script saved are now old enough to be genuinely surprising to me to look back on — which is really the whole point.
If you'd like a periodic nudge of your own, the Get Reminded page has the Bluesky and Mastodon bots, the daily email, and the RSS feeds — pick whichever fits how you already read.
If you have ideas on how the concepts of this site could be extended, let me know!
...those ideas might work their way into next year's summer vacation project.
Throughout the week, I attended keynotes and technical sessions, spent too much time walking around the exhibit hall, and had the chance to meet people working across quantum hardware, software, algorithms, HPC, and AI.
My parents and brother also tagged along for the trip, so outside the conference we had some time to explore Toronto together before eventually making our way to Niagara Falls after QCE.
Monday, September 14: Quantinuum and QUOPS
The conference started very strong.
The keynote that stood out the most to me during the week was given by Rajeeb Hazra, President and CEO of Quantinuum.
His talk focused on the next era of quantum error correction and how we should measure progress as the new systems move toward useful fault-tolerant quantum computing.
Quantinuum introducing QUOPS
One of the major announcements was the Quantum Universal Operation Performance System (QUOPS). QUOPS is a benchmark that does not depend on architecture, designed to measure how much useful quantum computation a system can perform and how quickly it can perform it. It was developed by Sandia National Laboratory, with input from Quantinuum and NVIDIA.
There was also an interesting connection that I had no idea about until the keynote. Dr. Hazra received both his master's and PhD in Computer Science from the College of William & Mary, which is very close to Old Dominion University.
Seeing someone who studied computer science basically down the road from ODU now runs one of the major quantum computing companies in the world was pretty inspiring.
The exhibit hall also opened on Monday night along with the posters. QCE had companies working on essentially every part of the quantum computing stack, from hardware and control systems to transpilers, software, and hybrid computing platforms.
Tuesday, September 15: IBM and the Road to Fault Tolerance
Tuesday had another keynote that I found interesting, although this one got technical very quickly.
Ali Javadi-Abhari, Head of Qiskit at IBM Research, talked about IBM's path from the noisy quantum computers we have today toward full fault-tolerant quantum computing.
One way he framed it that really stuck with me was as a complexity tradeoff between time (samples) and space (qubits).
Complexity tradeoffs behind quantum error mitigation and correction
On one end of the spectrum, error mitigation can work with fewer qubits, but you pay for it by repeatedly executing circuits and collecting many more samples. As you move toward full quantum error correction (QEC), you can reduce that sampling overhead, but now you pay in physical qubits and increasingly complex hardware.
I really liked this framing because fault tolerance is sometimes presented as a single jump from today's noisy machines to fully error-corrected quantum computers. In reality, it looks much more like a spectrum where different techniques make different trade-offs depending on what the hardware can support.
The other part that caught my attention was IBM's hardware roadmap.
IBM is building Quantum Starling at a new quantum data center in Poughkeepsie, New York, with the system planned for 2029. The architecture is highly modular, with multiple quantum systems working together instead of relying on one enormous (monster) processor.
The idea of connecting multiple systems also tied nicely into something that kept appearing throughout QCE: the future of quantum computing is probably going to be much more distributed and heterogeneous than a single QPU sitting by itself.
Wednesday, September 16: NVIDIA and the Quantum Supercomputer
Keynote by NVIDIA’s Krysta Svore
Wednesday continued with another keynote that really stood out to me; this time it was from Krysta Svore, Vice President of Applied Research for Quantum Computing at NVIDIA.
Her keynote focused on NVIDIA's vision for an accelerated quantum supercomputer, where CPUs, GPUs, and QPUs work together rather than as completely separate systems.
She also discussed CUDA-Q Logical, which is NVIDIA's freshly announced infrastructure for developing and testing fault-tolerant quantum applications, along with support for the QUOPS benchmark introduced earlier in the week.
I also had the chance to ask her for advice as someone coming from a computer science background.
I really liked seeing that connection between keynotes. On Monday, QUOPS was introduced as a way to think about useful quantum performance. On Tuesday, IBM went deep into how we might actually reach fault-tolerance. Then on Wednesday, NVIDIA was showing how fault-tolerant quantum computers could fit into a much larger software and accelerated computing stack.
The AI and quantum connection can roughly go in two directions. The first is AI for quantum, where LLMs are used to improve things like circuit generation, calibration, transpilation, error correction, or optimization. The other is quantum for AI, where quantum computers are used to accelerate or change how we perform machine learning itself. Right now, the first direction seems much more mature, but I think there is a lot of potential in both.
This whole AI-HPC-Quantum integration space is probably one of the research directions I am most excited about right now.
Thursday, September 17: AI Meets Distributed Quantum Optimization
Thursday had one of the talks that was most directly connected to my own research interests.
The work combines distributed QAOA with a generative model that produces candidate circuits for smaller optimization problems.
What made this interesting to me was that it combined almost every theme I had been hearing throughout the week: distributed quantum computing, GPUs, AI, optimization, and HPC.
It was also really cool seeing work from ORNL at QCE while I am currently doing an internship there.
By Thursday, it was apparent that hardware, AI, HPC, quantum algorithms, and software systems are no longer being discussed as completely separate areas. They are increasingly becoming parts of the same computing stack.
Friday, September 18: Microsoft and Presenting D-QEO
The last day of the conference started with the final keynote before it was finally my turn to present.
From scalar computing to quantum
I attended Matthias Troyer's morning keynote from Microsoft, and one idea from his talk really stood out to me: that quantum computing is simply the next computing paradigm. Over the last several decades, computing has moved from scalar to vector processing, massively parallel systems, GPUs, and now potentially to QPUs.
He also discussed Microsoft's Majorana 2 approach using topological qubits with the long-term goal of scaling the system to one million qubits on a single chip using fast, digitally controlled hardware.
Then, after spending several days watching everyone else present, it was finally my turn, and I was definitely nervous.
After sitting through a week of talks from people across the largest quantum companies, national laboratories, and universities, getting up to present our own research at QCE felt like a pretty big moment for me.
Presenting our D-QEO work
The main idea behind our work is different from the usual hybrid optimization approaches because we do not ask the quantum computer to directly find the final answer to a continuous optimization problem. Instead, we use it as a topographical preconditioner.
The quantum processor explores the discretized landscape and produces a probability distribution around the most promising regions of the search space before handing it back to the classical optimizer running on a GPU.
Rather than replacing the classical optimizer, the QPU helps it decide where it should search. For the separable functions studied in this work, we also decompose the problem into smaller independent subcircuits that can be evaluated separately. This allowed us to explore larger search spaces without requiring the entire problem to fit onto one very large quantum circuit.
What was funny was how well our work fit into the themes I had been hearing through QCE all week.
Again and again, people were talking about QPUs as accelerators inside larger classical workflows of CPUs and GPUs. Our approach follows a similar trend because we use the quantum processor for the part where it may provide something useful, and let the classical hardware handle what it already does well.
Presenting the work to a quantum-focused audience was also quite a bit different from presenting it to a more general computer science audience. The questions went immediately into things like circuit transpilation, hardware execution, decomposition, and scaling. I also had a chance to connect with other people afterwards for a potential collaboration.
Toronto and Niagara Falls
Niagara Falls from the Canada side
I obviously did not spend the entire week inside the convention center.
My parents from Hungary came to Toronto with me, which made this trip a little different from most of my conference travels so far.
Whenever I had some free time, we walked around Toronto and explored the city a little bit.
Since the conference was right downtown, it was easy to leave QCE and immediately be in the middle of the city.
After the conference ended, we also drove to Niagara Falls together for a day.
I had obviously seen Niagara Falls in pictures countless times before, but seeing it in person up close is completely different. It was a great way to finish the trip after spending an entire week thinking about quantum computing.
Looking Back
The CN Tower at night just outside of the conference venue
Overall, QCE 2026 was a really inspiring experience.
The original reason for going to Toronto was to present our paper, but just like with my trip to KDD, the conference ended up being much more than the presentation itself.
I heard from people building some of the most advanced quantum systems in the world, saw where companies are currently investing their efforts, learned about research that overlaps with my own work, and met a lot of interesting people from both industry and academia.
If there was one technical theme I took away from the week, it was definitely the growing convergence of quantum computing, AI, and HPC. I do not think these technologies are going to develop independently from one another, but rather in ways that support each other.
QCE made it feel increasingly likely that useful quantum computing will become one component of a much larger computing ecosystem. That also happens to be the part of quantum computing that I find the most exciting. Add in presenting our own work at a conference I was honestly nervous to present at, meeting people from across the field, exploring Toronto with my parents and brother, and finishing the trip at Niagara Falls, and it was a pretty memorable week.
I would also like to thank Dr. Nikos Chrisochoides and IEEE Computer Society for making this trip possible.
Yesterday Donald Trump ordered the Interior Department to change the name of Lake Ontario to “Lake America” in its Geographic Names System (GNIS). If your first thought on hearing the news was “I wonder if library subject headings will change?” you’re probably a librarian. Or at least someone who shares some of my uncommon interests.
Mind you, just because a government changes its name for something doesn’t mean that anyone else outside of government will go along with the change. But historically, many organizations have relied on the GNIS for names to use in maps, location designations, and place references in general. The Library of Congress is one such organization, and when in 2025 Trump made similar orders for the Interior Department to rename the Gulf of Mexico and Denali, bypassing the usual professional and stakeholder review for such changes, the Library soon made correspondingchanges to its name and subject headings. Other American libraries went along with it, adopting the new names as well (often by default, as they automatically imported catalog records from shared systems and commercial providers).
So it’s quite possible that we’ll soon start seeing headings like “America, Lake of (N.Y. and Ont.)” in catalog records at the Library of Congress and elsewhere. At this point, though, I don’t plan to make such changes in the Online Books Page catalog. To understand why not, it’s worth reviewing why libraries choose the headings we use in our catalogs. As I see it, there are three main concerns our heading choices should satisfy:
Findability: The main purpose of headings in a library catalog is to help people find the information resources they need. To do that, the headings need to employ language our readers are likely to use in their searches, and that they recognize and follow in their browsing. That’s a strong argument for using the same terminology that people use in their everyday speech and writing, even if that differs from official terminology.
Respect: It’s also important that the terminology we use in our catalog respects both our users and the people associated with whatever we’re naming. For instance, if someone announces a change in their own name, or the name of an organization they run, we typically respect their autonomy by reflecting that change in our catalog, even before the name becomes commonly used by others. Likewise, if a group of people associated with a thing or a name say that the name we use is insulting, or would be better changed to something else, we listen to them, and then often change the name if there aren’t compelling reasons not to.
Standardization: This is an often under-appreciated reason libraries choose the headings they do: Not only does it help users find things if we use the same names and headings consistently across our library systems, but it’s also often much easier for us to maintain our own catalogs if we use the same names and headings as our peer libraries do. That lets us easily exchange records, share cataloging work, and conduct searches across multiple libraries. Few libraries have the resources to create and maintain their own system of headings, but since the Library of Congress does, and is a major source of catalog records for libraries, most American libraries adopt their headings rather than maintain their own terminologies. The advantages of standardization typically outweigh concerns about less-than-ideal heading choices, or the differences in the needs of the Library of Congress’s primary user constituency (which ultimately is the US Congress) and the needs of local library communities.
Given these concerns, changing Lake Ontario’s name in my catalog would not be in the interests of the library I run, or of its users. Changing its name to Lake America would make books I list about the lake less findable, because few people will normally use that name. I’m not aware of anyone besides Trump using or calling for that name in the past, and the lake’s been continuously called “Ontario” since before either the United States or Canada were countries. (“Lake America” might get picked up by this president’s enthusiastic fans, his sycophants, or people and agencies he can order around, but those represent only a minority of my users.) The change is also blatantly disrespectful, not only because it ignores the usage and preferences of the people who live around the lake, but also because Trump announced the name change, for a body of water shared by the United States and Canada, explicitly to retaliate against Canada in a trade war that remains ongoing.
That only leaves standardization as a reason to go along with the change, should the Library of Congress adopt it (which I should note is not yet certain). But with the development of linked data and related technologies in libraries, and with software I’ve developed for The Online Books Page, it’s no longer as important as it once was that my library use all the same headings as other libraries. Simply holding back certain revisions to a standard terminology is easier to sustain than other deviations I might make from that terminology. I’m already doing some automated revision of headings in records I import to our extended shelves, and it’s not a big change to make those automated revisions use older headings for certain names and subjects. I’ve also had software support for automatic conversions between local library headings and standard Library of Congress headings for years. I can use that to seamlessly support links between searches in my catalog and searches in catalogs that use Library of Congress headings. (If I know other libraries are also making similar changes, I can also maintain data to support seamless mappings to their catalogs.)
So whatever the Library of Congress or other libraries do, Lake Ontario will stay Lake Ontario in the Online Books catalog. Or at least it will as long as that remains the predominant everyday name for the lake, and there hasn’t been a properly consultative process to change it. I will be making some changes in my systems and policies to accommodate this decision. I can elaborate more on how I’m doing that in a followup post if readers are interested.
The same day I posted Downgrades Peter Oppenheimer et al from Goldman Sachs (GS) weighed in on the same topic with Competition for Capital:
"There are two themes dominating our investor conversations: the impact of AI and the rise in interest rates. The two issues are linked as demand for capital from both the private and public sectors increase. In the private sector a surge in capex spending to fund AI infrastructure has eaten into free cash flow and forced companies to raise more in debt and equity markets. Meanwhile, government borrowing needs have increased as priorities shift towards upgrading critical infrastructure, energy security and defense at a time when cyclical inflationary pressures driven by higher energy prices are also resulting in higher policy rates. The combination has pushed up the cost of capital. As recently as 2022, for example, 30-year bond yields in Germany and Japan were close to zero ... A rise in yields, together with more uncertainty (over geopolitics and the future impact of AI) have, in combination, pushed up the cost of capital."
As one would expect from the professionals there is a lot to digest in their report. Below the fold I discuss the details.
GS are interested in capital allocation, and the the comparison between stocks and bonds. So they focus on the Equity Risk Premium (ERP), the extra you are paid for investing in stocks over bonds because the risk is greater:
This rise in yields comes after a near-record period of equity outperformance relative to bonds over 10-year holding periods
As equities have continued to outperform bonds, equity risk premia have fallen back to levels last seen in the late 1990s, leaving equity markets more vulnerable to further increases in bond yields. That said, in the US and Japan, the ERP has bounced off recent lows
Remember what was happening in the stock market in the late 90s?
Just as it did then, everything in the stock market just fine:
But despite the rise in yields, nominal GDP and profit growth remain strong and, until recently, the impact of rising bond yields has been offset by the strength in corporate profits. Earnings growth has been the main driver of equity returns over the past 18 months in all regions (Exhibit 5). This strength has meant that PE multiples have either been flat (Japan and Europe) or have fallen (the US, Asia and EM).
In the US, in particular, the forward PE multiple for the S&P has come down from 22x at the start of the year to 19x, in line with its long-run average, despite the market being close to its all-time high.
Note that the standard explanation for the valuation of a stock is that is the Net Present Value (NPV) of the company's future earnings (or more strictly its future free cash flows). An increase in interest rates decreases the NPV of these earnings, and thus decreases the stock's PE.
GS flags this:
The combination of higher cost of capital and greater capital intensity in the tech sector has reduced the value of their future cash flows and triggered a de-rating in the biggest companies. The dominant capex hyperscalers in the US now have a forward PE close to the average of the rest of the market.
A very large chunk of the S&P is the hyperscalers, whose future free cash flow has vanished.
GS describes the growth of hyperscaler borrowing:
While technology profit growth has remained strong, the surge in capex spending among the hyperscalers has increasingly eaten through their free cash flow (Exhibit 8) prompting companies to look for alternative sources of funding, in the credit and equity markets. The capex spending growth for AA-rated issuers has been substantial: the 65% year-over-year growth in Q2 marks the 10th consecutive quarter that aggregate AA capex growth exceeds 35%. They have also turned to the convertible bond market where volume has also increased year to date reaching $135 billion in the US, with AI-related borrowers driving 44% of total issuance, on our estimates.
Supportive profit growth has been accompanied by positive earnings revisions, with 2026 and 2027 estimates moving higher across major regions. There are four broad areas which have been driving much of this profit growth. First, technology earnings continue to be very robust. Second, rising energy prices have pushed up profits in the commodity sector. Third, banks have generally enjoyed strong earnings backed by positive nominal GDP growth, steep yield curves and strong private sector balance sheets. Finally, sectors such as Industrials have benefited from the surge in AI capex spending which has spilled over into improved revenues for the ‘pick and shovels’ of AI infrastructure. The breadth of the earnings' growth has also increased the opportunity for investors to diversify across sectors as well as countries.
Exhibit 11
The thing that worries GS is whether these current earnings, especially those of the hyperscalers are sustainable:
Bubbles in earnings, rather than valuations, have occurred in previous periods. The Bank sector, for example, briefly became the biggest sector in the S&P 500 in the run up to the financial crisis of 2008/09 (Exhibit 11). Unlike the technology bubble of the late 1990s, or the Japanese bubble of the late 1980s, Bank stocks did not experience a major valuation bubble at the time, but the surge in earnings that drove the sector's outperformance proved to be unsustainable.
The Paris-based Organisation for Economic Co-operation and Development (OECD), the International Monetary Fund (IMF) and the International Institute of Finance (IIF), the voice of global banking, on Wednesday highlighted the dangers of soaring interest rates on $365tn (£275tn) in global borrowing.
In its quarterly debt monitor, the IIF predicted a “structurally debt-intensive future” as governments and companies scramble to invest in new technologies and bear the costs of ageing societies.
“The buildup in global debt is set to accelerate as governments and corporates compete to boost growth and secure their positions in an economy reshaped by structural changes,” it said.
On September 10, 2026, Jacqueline J. Dearborn, Library Innovation Lab Fellow, gave a virtual talk about the Public Data Project’s International Data Infrastructure Research and Development effort. The transcript below has been lightly edited for clarity.
Hello everyone, before I get started, I want to give a shout-out to Molly Hardy and Jack Cushman at the Harvard Library Innovation Lab for giving me this opportunity to reflect on some of the most pressing challenges of our time.
With Jennifer Chapman as my counterpart, we’ve spearheaded the Public Data Infrastructure Research and Development project and through this work, we’ve gotten the chance to zoom way, way, way out and take a look at the global public data landscape.
Obviously this is a very huge, thorny, and complex topic. Much larger than any one talk could possibly resolve. So today, I share some preliminary insights that I hope will help us think together about public data infrastructure, long-term stewardship, and collective responsibility. In this spirit, I hope this talk is less of an ending and more of a beginning for all of us that spurs many future conversations to come.
Thriving. Struggling. Gone. Reborn?
Since I know some of you but not all of you, I thought it might be useful to tell you a little bit about myself. So, I began my library career right here at Harvard in 2011. I worked as a grad student at the Museum of Comparative Zoology’s Ernst Mayr Library to process and digitize collections for the Biodiversity Heritage Library. Since then, I’ve spent most of my time working on the infrastructures that hold public data.
As you can see by my career timeline in logos, the institutions and platforms have changed, but the mission to open up knowledge for all really hasn’t. I’ve moved through government, academia, libraries, museums, nonprofits, and research organizations. I’ve followed my purpose and, perhaps less romantically, I have also followed a lot of three-to-five-year funding cycles.
So while many of you may know me best as a member of the BHL Secretariat and Data Manager, I’ve also worked on and contributed to a wide range of other infrastructures, including Internet Archive, Digital Commonwealth, ADS, Boston Open Data (now known as Analyze Boston), Resource Watch, Dataverse, MassDocs, and a myriad of Wikimedia projects.
I suppose the charitable word to describe my career trajectory would be tenacity — or maybe just a stubborn refusal to abandon the work that I believe in. And all this moving around has given me an unusual vantage point. I’ve spearheaded new initiatives and I’ve been around projects that are thriving. I’ve watched others struggle, and heartbreakingly, I’ve watched some go extinct altogether. I’ve also seen some reborn and flourish again.
And when you’re working on the inside, you tend to measure success by fairly standardized metrics: Did we digitize the collection? How many pages? Did we build the API? Did we transcribe it? Did researchers use it? How many? All of these metrics definitely matter. But lately, I’ve become interested in different measures of success because data generation and stewardship is so incredibly labor-intensive.
Measuring success beyond funding cycles
Lately, I’ve begun to ask: Will the data still exist 20 years from now? What about 50? Or 100? And I don’t simply mean, will the files still be somewhere. I mean: Will anyone still understand what the data is, why it exists, and how to use it? Will there still be a community that cares for it? Will there still be enough funding and expertise to repair and maintain it?
These are very different metrics and they entail much more than preserving just the bits and bytes. And honestly, this is part of why I had to take the last year off. For me, this has been a critical moment to reflect. I spent a long time being very close to the machinery, especially in my last role as a data manager at the Smithsonian. And now I’m reflecting on the larger system that this machinery sits inside of.
And since I’m still figuring things out, I’m not going to give you some grand theory of the universe today. But I will give some personal observations I’ve made over the years. The things I’ve seen. The things I keep coming back to. As well as some uncomfortable truths that I think are becoming harder for us to ignore.
Straw, wood, or brick?
So before we get too serious, I want to begin this talk with a little story called “The Three Little Pigs.” This may not be where you expected things to go, but please bear with me. So, we’ve got 3 little pigs. One builds a house made of straw. One builds a house made of sticks. And one — the boring, sensible pig — builds a house made of brick. And then, of course, the big bad wolf comes along, huffing and puffing. And, well, we all know how this story ends.
I’ve been thinking about that story lately as a metaphor, and how it maps rather uncomfortably well onto the infrastructure that holds public data. Because we have built an awful lot of houses with data in them. Some are made of very strong materials. And some are built with straw.
And some infrastructure is beautiful, expensive, and sophisticated. But, nevertheless, it depends on one or two people holding the keys. And for a very long time, we as data home builders have tended to ask questions like: Is the data safe? Did we back it up? Is it online? Can people find it? Can the machines read it?
All of those are great questions. But I think we’re increasingly having to ask the harder ones: Is the house itself structurally sound? Have we built on shaky foundations? What happens if the wolf comes around? Are we ready? Because the big bad wolf isn’t really one thing, is it? Sometimes it’s a change in institutional priorities. Sometimes it’s a funding cut. Sometimes it’s an aging server well past its warranty. Sometimes it’s a person leaving after twenty-five years. And sometimes … it’s simply just time.
And that brings me to the loaded word I’ve been dancing around since the beginning of this talk: infrastructure.
The houses that public data live in
So far we’ve basically defined infrastructure as “some houses with data in them.” This word “infrastructure” may conjure thoughts of code, servers, platforms, networks, storage, software. But I think we’ve inherited a rather narrow technical definition of infrastructure. Because if the goal is to preserve and transmit knowledge across generations, then infrastructure is actually the entire system that allows data to be created, interpreted, trusted, preserved, and shared in the first place.
So let’s map it out. What are some of the core components of infrastructure? Well, there is obviously technology and data but we also have people, expertise, funding. And increasingly, I’ve been thinking about governance and the law as the wrapper around the entire house where human roles, responsibilities, decision rights, and accountability mechanisms all define how those other pieces actually work together.
Sure, there are twenty-seven ways to make this mental model more academically rigorous, but for now, I’m finding it really useful. Today I want to discuss the parts of the house that I know best which are the technical and deeply human dimensions. Because the various components in the house are not independent. They’re interdependent: Technology without expertise doesn’t get us very far. Expertise without funding disappears. Funding without governance becomes chaos. And so on.
So we shouldn’t think of infrastructure as a handful of separate pieces rather we should think of it as the house they build together. Like the electrical, plumbing, HVAC, and so on. All of these things have to work together if you want the house to function and keep its contents safe. And once you look at infrastructure this way, some of the things happening right now start to look less like isolated crises … and more like maybe we have a lot of straw houses and in them is our incredibly valuable public data.
When the house becomes brittle
Katherine Skinner, Director of Programs at Invest in Open Infrastructure, cuts straight to the issue for us:
Data is disappearing because the infrastructure in which it is nested was already brittle.
And in real time, we are watching that brittleness become increasingly visible. I’m not trying to catastrophize here, because this isn’t one single big-bad-wolf moment isolated to the U.S. Everywhere, infrastructure providers are facing some combination of AI bot traffic, leadership and funding shifts, aging technology, and the loss of key personnel.
Infrastructure survival can depend on a remarkably particular set of circumstances: a grant; a director; a willing host institution; a community champion; sometimes just a handful of people who really cared. And one of the most persistent “wolves” I have noticed has been chronic underfunding.
There is evidence for this. Invest in Open Infrastructure examined more than $550 million in grant funding across the public data ecosystem. Of the funding that flowed directly to infrastructure, nearly two-thirds supported research, development, and innovation. But only about 20 percent supported core operations and maintenance. So we are very good at funding shiny new things, but we are not so good at funding them to endure.
And when operational funding begins to evaporate, the effects compound. People leave; institutional knowledge leaves with them; the system stops evolving; users lose confidence; and the community shrinks. And from the outside, it might look as though the data is still there. And sometimes the files are technically still there. But the community-driven ecosystem that made new data generation possible begins to fall away.
When infrastructure stops evolving, in a sense, it stops living. And that is a different kind of loss. For those of us who have been close to the machinery, fighting to keep these systems alive, that loss can feel a lot like intense grief. Which brings me to an important distinction. When public data is endangered, rescuing the bits and bytes may be absolutely necessary. But rescue is only the first leg of a much longer intervention.
Rescue preserves the data. But stewardship preserves the conditions that allow that data to remain usable, meaningful, connected, and dynamically alive. If we copy the files somewhere safe but separate them from the communities, provenance, expertise, and infrastructure that gave them meaning, have we really secured them for the long horizon? Perhaps rescue is the emergency intervention. Stewardship is what has to come next.
Data, information, knowledge
So bear with me while I take you back to library school for a moment: A geo-coordinate can be recorded. A page image can be preserved. A database record can be copied. And we have the LOCKSS principles to safeguard this data. We can make multiple copies. We can checksum things. We can put them in geographically distributed storage. And we can keep them in highly durable formats.
But saving those data points doesn’t really mean we’ve preserved how to interpret and use them. And this difference really matters. Data can be recorded and replicated; information gives it context; knowledge tells us what it means and what to do with it — and much of that knowledge does not live in the files themselves. It lives in people. And that creates a major preservation problem. Because we are very good at preserving data. But people? You can’t LOCKSS a person … at least not yet.
You cannot LOCKSS a person
Since antiquity, human beings have always transmitted knowledge. We spoke it aloud; carved it into stone; wrote it on papyrus and parchment; printed it into books; and now we replicate it as digital data across the planet. The medium keeps changing, but throughout history, knowledge has survived through a continuous chain of human stewardship.
So despite all our new-fangled technology, I’m here to tell you something: the oral tradition is alive and well, folks. It never actually went away. Because ultimately, all of this knowledge is for human beings. We created this data. We are the core users of the public data ecosystem. And I think it’s time we start prioritizing the human use case — not just the AI use case.
AI can dramatically accelerate our ability to discover, connect, and use shared knowledge. But it does not eliminate the need for human stewardship. Its greatest promise is to expand humanity’s capacity to learn, understand, and act. And of everything in our infrastructure house — technology, funding, community, governance, law, expertise — the part I think we understand least well is the deeply embedded human one.
The people. Us.
Because we have a lot of good ideas about preserving data, but you can’t LOCKSS a person. You cannot make a second copy of somebody’s twenty-five years of experience and put it in another geographic location. You can’t checksum institutional memory. This kind of knowledge exists everywhere. We often call it tacit knowledge.
Sometimes that just means knowledge that has never been written down because, quite literally, its only storage medium is a human being. I’ve watched this happen. And I’ve participated in it too. I’m guilty: I’ve been the person who knew where something was; I’ve inherited systems from someone who later retired; and I’ve watched people leave organizations carrying extraordinary amounts of infrastructure out the door with them.
That’s changed how I think about resilience. Because sometimes the first failure isn’t a funding cut or a downed server. It’s a person. Somebody walks out the door, and the cascade begins. So when we talk about investing in infrastructure, we need to get much more comfortable saying something that sounds obvious, but apparently is not:
People are infrastructure. And if we don’t invest in the people who maintain, interpret, govern, and evolve our public data infrastructures, then I can’t really see how public data survives this decade. And that brings us back to the house. Because the same problem that happens when we put too much knowledge in one person’s head also happens when we put too much of our technology in one bespoke system.
We don’t need a gajillion Ferraris
So, the year is 2026 and we have built a remarkable number of beautiful, one-of-a-kind data platforms. Some of them are genuinely brilliant. They solve very specialized problems. They have really cool features and slick interactive interfaces that are optimized for a particular collection or community. And that’s not necessarily a bad thing.
But we’ve also developed a habit of building things that only one or two people in the world know how to maintain. And all that sparkle may be masking some very shaky foundations. And at some point recently I started thinking of data platforms as custom Ferraris. They’re beautiful, impressive, rare; but if something goes wrong, you really gotta hope the mechanic is answering their phone and won’t ask you for a hefty ransom to fix it.
We need a shared FOSS ecosystem
You know what I think? I think we need more Honda Civics. Not because Hondas are better cars, but because there is an enormous ecosystem around them. There are parts, mechanics, manuals, standards. There are many people who can look under the hood and have some idea what they’re looking at. That’s a really important property for public data infrastructure to have.
Because the question shouldn’t just be: can we build it? The question should also be: can more than one human maintain it? Can someone else easily contribute to its code base? Can another institution host it? Can a new person learn it? Can the feature that is genuinely unique sit on top of something shared? And I think this is where open source infrastructure becomes much more compelling than simply “free software.” The real value is not just that the code is available. The real value is that we can distribute the capacity to understand and maintain it.
We don’t need every institution to build its own bridge to some siloed data island. We need to build common infrastructure across datasets and collections using free and open source software, and then we do our distinctive, special-snowflake things on top. Because we need more houses built with brick. And in the virtual world, I think one of the key ingredients for a strong house is Free and Open Source Software (FOSS).
When convenience costs us sovereignty
Unfortunately, there is a rather enormous emerging issue with everything I just said. The FOSS ecosystem does not maintain itself. People maintain it. And many of the institutions that we have come to depend upon are losing the internal capacity to build, deploy, operate, and maintain open infrastructure.
As universities and research organizations outsource more of their technical capacity amidst constrained budgets, the immediate convenience is obvious: someone else manages the servers; someone else handles the updates; someone else keeps the lights blinking. And over time, institutions can stop employing the people who know how to build and operate infrastructure independently.
And this strategy might save us a short buck — but does it get us to the long horizon? One of my mentors, Professor Peter Cornwell, put this much more directly:
FOSS system platforms have advanced functionality; they share vigilance and overcome emerging security hazards via an expert community; they don’t attract license fees and operate indefinitely.
Many institutions have lost these skills: they’ve become support organizations for proprietary products. We must urgently set about restoring these capabilities.
And I think that last point is the one I want to underscore. We may eventually still have the FOSS code. But will we still have enough people who know how to deploy it? That is why there is some urgency here. If we want to preserve the option of building a global data commons on shared, open, interoperable foundations, we have to invest in the people capable of doing that — and we need to do it now.
Otherwise, we may discover that we still possess the blueprints … but have lost access to all of the builders.
What does resilience actually look like?
So I’ve just painted an elaborate picture using the Three Little Pigs and custom Ferraris to illustrate what human and technical fragility look like from the inside. Which brings us to the central question: What does resilience actually look like?
And for this, I want to borrow from biology — because I’ve spent most of my career in biodiversity and I can’t resist. Healthy ecosystems aren’t resilient because every component is invulnerable. They’re resilient because they are diverse, distributed, redundant, interconnected, adaptive, and regenerative. Instead of monocultures, we have biodiversity. Mother Nature really is the best designer, isn’t she?
And I think public knowledge infrastructure needs to start thinking in these terms. It’s not: How do we make one organization, brand, or institution the end-all, be-all house for the data? It’s: How do we build a system capable of surviving the failure of any one part?
That principle has to apply across the entire infrastructure house: Shared technology. Distributed expertise. Inclusive communities. Diversified funding. Transferable legal arrangements. Adaptive governance. Not one institution holding all the cards. Because if an organization has six months of runway left, that’s not resilience. That’s a countdown.
Resilience doesn’t mean making every component perfect. It means creating a system that is less dependent on perfection. We’re not trying to build one house that the wolf can never blow down. We’re trying to build an ecosystem where losing one house is a recoverable event — where the wider community has the resources, rights, knowledge, and capacity to carry the work forward.
We aren’t starting from scratch
Now, I realize I have spent quite a bit of time telling you what is going wrong. But I want to be very clear about something. We are not starting from scratch. In fact, many of the ideas we need already exist: We have LOCKSS and the wonderful idea that Lots of Copies Keep Stuff Safe. We have trusted repositories. We have open-source platforms. We have persistent identifiers. We have interoperable standards. We have a global community of experts operating all over the world.
And we already have enormous networks of libraries, archives, museums, universities, research organizations, governments, nonprofits, community groups, developers, collectives, and individual experts who care very deeply about keeping public data both sovereign and alive.
So the problem is not that humanity has no idea how to do this. Because we have many of the pieces. What we have not yet done is connect those pieces into a sufficiently resilient global system. LOCKSS taught us something incredibly important about data: Don’t depend on one copy. I think the challenge before us now is to take our learnings from LOCKSS one step further and apply them to not just the data but all of the dimensions of infrastructure:
Don’t depend on one institution. Don’t depend on one funder. Don’t depend on one platform. Don’t depend on one country. Don’t depend on one technical team. And certainly don’t depend on one person.
A global mycelium network
And there is something that makes a distributed system like this possible. Standards are the agreements systems make with one another. Law and governance are the agreements people and institutions make with one another. Together, these form a global mycelium network. Persistent identifiers. Shared protocols. Common data structures. Open-source stacks. And open governance processes that determine how we work best together.
Yes, these things feel painfully boring. Nobody is going to give a keynote about the metadata field they viciously argued about internally about for five years and finally standardized in 2017. But this is what makes interoperability and continuity possible. Shared agreements are what allow data and resources to move between systems. Because ultimately, standards and governance are relationships and social agreements expressed technically, organizationally, and legally.
And that’s why the next paradigm won’t be a collection of isolated successful data projects. It will be a unified community with shared foundations to steward the commons. And from my technical point of view, a shared FOSS ecosystem isn’t a nice-to-have. It’s the missing requirement for long-term public data stewardship. Which brings me to my next thought.
Towards a global data commonwealth
What would it look like if we stopped thinking about public data infrastructure as something that individual institutions steward and are ultimately responsible for? What if we thought about public data and infrastructure as a kind of shared inheritance? Not owned by one library or one institution or one country or even one generation. Something we steward collectively because the consequences of losing it are collective.
I don’t have a fully formed answer for what a global governance model looks like for that. But I’m hoping some of my friends here at Harvard Law Library and the people in the audience can lend a helping hand. And I actually think it’s useful to acknowledge that I don’t personally have all the answers. Because none of us do. And it is only through our own humility and our ability to cooperate and collaborate that any of this actually works.
I don’t think the answer lies in some bold visionary coming onto the scene to save the day. Or creating another organization, consortium, or intergovernmental interface. We already have all those things. Instead we need to answer the harder questions here: How do we connect what already exists? How do we change how resources flow? How do we protect and distribute expertise? How do we prevent single points of failure across all dimensions?
So as we near the end of this talk, I think we’ve evolved the narrow technical definition of “infrastructure.” To me, it’s beginning to sound a lot less like a bunch of data platforms and a whole lot more like a commons. Or perhaps a global commonwealth? Imagine a truly global network with different roles, different skills, and different interests — bound by the shared mission to steward public data and carry it forward for the collective benefit of all and future generations to come.
The future of the commons
A good friend of mine, Ben Vershbow, is the founder of As We May Think, which is a new lab dedicated to strengthening the digital commons. When I asked him what the future might look like, he had some interesting insights. He said:
A commons concentrated in a few prestigious and powerful institutions is a fragile commons. The future lies in a federation of diverse communities — attuned to their contexts, connected through shared standards and infrastructures — capable of stewarding knowledge together.
And I think the important word here is diverse. Not diversity as decoration or inclusion as a slogan. Diversity as resilience. A monoculture is fragile. A resilient commons needs many kinds of stewards: people, collectives, communities, organizations, and institutions — across geographies, disciplines, cultures, and various levels of power.
And if we genuinely want distributed stewardship, then we need to distribute resources more effectively and efficiently: money; authority; expertise; opportunity. Because we cannot decentralize stewardship if we are centralizing all the resources.
What now?
I’ve spoken to you today from the parts of the infrastructure house that I know best — the technical and intimately human side. But I’ve only lightly touched on the legal and international governance dimensions that will determine what a resilient global data commons might look like. That’s where my counterpart Jenn and the brilliant folks at the Library Innovation Lab, and many of you here, come in.
Jenn frames this moment as both a crisis and an opportunity to redesign and rethink the current paradigm. In her forthcoming paper, “Responding to the Global Access to Information Crisis through Sustainable Information Access Goals,” she writes:
In the face of mounting challenges to global information access, it is imperative that librarians, information professionals, and other stakeholders develop resilient, connected collections that counter current disruptions and withstand future attacks.
And I think that brings us back to the central point. None of us will solve this alone. The challenge is too big for any one institution, any one discipline, and probably too big for any one country. We may not know exactly what that future looks like yet. But I think we’ve reached the point where we need to start designing for it.
So who will steward it?
So, back to our original question: Who will steward humanity’s public data? I suppose I should admit something. I don’t actually know. But perhaps the answer starts with the people in this room. Perhaps many of us came here today because something about this question stirs something in your gut that transcends your career, your home institution, grant obligations, or personal loyalties to any particular brand, project, or platform.
And if we’re really serious about saving public data, then I think that asks that we all reflect a little bit on some very difficult questions:
Have we built systems that depend too much on our own expertise, authority, institution, or vision?
Are we making it easier for others to carry the work forward — or harder?
And perhaps most uncomfortably: are there moments when stewardship means holding on, and others when it means letting go?
As we know, the public data landscape is shifting and we are experiencing a lot of data losses. Now more than ever, our community needs to come together, globally. To do this work, it will require that we think way beyond ourselves and commit multiple acts of humility, generosity, and service. Because ultimately stewardship means building for a future that does not depend on us.
Thank you all so very much for letting me have this soapbox moment. I’m incredibly grateful to be in a room filled with brilliant people willing to think about the future of public data — and I hope to collaborate with some of you to get busy building brick houses for all the data we’ve been rescuing as of late.
The device to the right is the external CD RW backup, which I purchased
from Ebay, since I (foolishly) discarded John’s when we were clearing
out his apartment. I need this player in order to be able to be able to
read the CD backups of his music that he left, since the TASCAM 788
wrote the backups in a proprietary format.
Many years ago, I wrote a tips on job applications post, because I had been seeing a lot of poorly formatted and written applications. While good formatting and tailoring to job postings are important, the content is obviously important as well. Again, there is so much literature out there, but what I’ve seen recently has … Continue reading "Yet another resume (cv) writing guide"
The $28.5T forecast compares to U.S. Q1 2026 nominal GDP of nearly $32T, with the estimate for the market of AI enterprise applications of $22.7T about 70% of total U.S. economic output.
Sam Altman and Dario Amodei just aren't this good, but their projections of their Total Available Market (TAM) are still turning out to be vastly optimistic. In AI's Affordability Crisis I showed evidence that the AI platforms could no longer afford the massive subsidies they were using to artifically inflate demand for their product, and that reducing the subsidies had made their enterprise customers reconsider their enthusiasm for deploying them. This is leading to investors belatedly realizing that AI platforms' projections of their TAM and thus their valuations are totally implausible.
This re-calibration is just one of the many signs that the AI bubble is about to deflate. Below the fold I present a necessarily incomplete list of them, which I will try to update as more appear.
AI now accounts for nearly half of all IG issuance, 87% of VC funding and a growing share of HY, underscoring how deeply the AI investment cycle has penetrated every corner of finance.
Here’s the title page of this month’s Panmure Liberum market update from strategists Joachim Klement and Francisca Reis. Our emphasis in bold below:
In 1929, the cyclically-adjusted P/E-ratio (CAPE) of the S&P 500 reached 32.6x according to Prof. Robert Shiller’s data. This was 1.8 standard deviations above trend at the time. In 2000, the CAPE reached 44.2x, or 3.3 standard deviations above trend – a clear sign of a bubble. However, as our chart below shows, earnings in both instances were within normal range, less than one standard deviation above trend.
Today, the CAPE is at 41.0x, or 2.9 standard deviations above trend. Once again, we are clearly in bubble territory for stock market valuations. However, unlike in previous bubbles, we are having extremely high CAPE at a time when earnings themselves are 1.8 standard deviations above trend. In other words, we are in a valuation bubble at a time when earnings are in a bubble themselves.
If we correct for the earnings bubble, the current CAPE would be 67.6x or 4.6 standard deviations above trend, a bubble that surpasses anything ever seen in US history by an extreme margin. If valuations followed a normal distribution (which they don’t, so don’t take this literally), this would happen in 0.00019% of months or once every 43,432 years.
A massive chunk of this quarter’s blockbuster “growth” didn’t come from selling more software, shipping more microchips, or delivering more packages. Instead, it came from an accounting rule that forced massive, illiquid “paper gains” onto the income statements of tech giants.
In Q1 2026 alone, just three companies—Alphabet, Amazon, and Nvidia—reported a staggering $69.2 billion in non-operating windfall under their Other Income and Expenses (OI&E) lines. When you run the macro numbers, this single accounting phenomenon artificially inflated the entire S&P 500’s quarterly earnings by about 12%.
That 12% takes the factor from 67.6 to 75.7. Wang notes that correcting for the 12%, "the Q1 2026 earnings growth rate will not be very different from the 5-year average of 16%." In other words, the bubble is feeding upon itself — increased stock prices causes increased earnings causes increased stock price ... But suppose, for example, that OpenAI were to suffer a down round. Then Alphabet, Amazon, Nvidia and others who included paper gains in the "other income" on the way up would have to include paper losses in their income on the way down, amplifying the crash. OpenAI's last round valued the company around $750B and they were planning an IPO for at least $1T, but had to postpone it.
The foundation model era — roughly 2020 to 2025 — is over. The forces that defined it have inverted. Open source models have reached frontier performance while inference costs approach zero, exposing what was always structurally true: pre-training large language models at scale is not a durable competitive moat. The US government's formal designation of Anthropic as a supply chain risk in February 2026 accelerated a transition already underway — but did not cause it. The paper argues that the AI industry is restructuring simultaneously along four axes: economic, as the circular financing structure that inflated foundation model valuations collapses; technical, as the pre-training scaling paradigm gives way to post-training optimization, test-time compute, and agentic composition; commercial, as application-layer integrators displace the foundation model companies whose commodity they now consume; and political, as the government asserts its historic role as gatekeeper of strategic technology. These are not separate disruptions. They are one structural shift, arriving together.
The most consequential and least-discussed dimension is the permanent divergence between commercial AI and a classified national security AI track — built on different data, governed by different rules, and developing capabilities the public ecosystem cannot see, measure, or govern. Like every dual-use technology that has altered the calculus of state power, AI is being brought under government authority not by design but by the structural logic of what it is. The paper further argues that open-weight models are the counterintuitive instrument of sovereign control: a government that holds the weights commands the capability on its own terms, without dependence on vendor policy, financial continuity, or personnel clearance. The apparent openness of distributed model weights is, from a deploying government's sovereignty standpoint, the most governable architecture — because what cannot be withdrawn by a vendor's API policy cannot be taken away.
Grogran implicitly assumes that AI is useful but that the current margins and thus the valuations aren't sustainable.
DeepSeek-V4-Pro is priced through its API at $1.74 USD per 1 million input tokens on a cache miss and $3.48 per million output tokens.
That puts a simple one-million-input, one-million-output comparison at $5.22. With cached input, the input price drops to $0.145 per million tokens, bringing that same blended comparison down to $3.625.
That is dramatically cheaper than the current premium pricing from OpenAI and Anthropic. GPT-5.5 is priced at $5.00 per million input tokens and $30.00 per million output tokens, for a combined $35.00 in the same simple comparison.
Claude Opus 4.7 is priced at $5.00 input and $25.00 output, for a combined $30.00.
Six times cheaper for an equivalent product is likely to cause pricing pressure on the incumbents, who need to raise not reduce prices. DeepSeek is likely also subsidizing usage, but they do have real advantages. First, they have fewer resources so are forced to to be inventive. Second, the 40% of infrastructure capex that isn't the racks is much cheaper in China. Third, the power component of opex is much cheaper and more available in China. The result is:
In practical terms, DeepSeek does not need to win every leaderboard row to matter. If it can deliver near-frontier performance on many enterprise-relevant agent and reasoning tasks at roughly one-sixth to one-seventh the standard API cost of GPT-5.5 or Claude Opus 4.7, it still forces a major rethink of the economics of advanced AI deployment.
DeepSeek-V4-Pro-Max is clearly the strongest open-weight model in the field right now, and it is unusually close to frontier closed systems on several practical benchmarks.
While GPT-5.5 and Claude Opus 4.7 still retain the lead in most direct head-to-head comparisons across the company's benchmark charts, DeepSeek V4 Pro gets close while being dramatically cheaper and openly available.
Dan Davies agrees that the margins aren't sustainable but differs as to why in tokenalysis and john henry:
Which then brings to mind another issue – how confident are we in the pricing power that underpins that 60% gross margin in the first place? In the last paragraph I was talking about the R&D equivalent of a price war, but the normal kind is also possible. The combination of price-sensitive B2B customers, big fixed costs and rewards going to the dominant player doesn’t suggest to me that pricing power is going to be sustainable indefinitely.
But, I think there’s a danger of missing the big picture here. Which is that, when large companies are telling their employees to be sensible and use AI tokens wisely, then the game is up. The race is over and John Henry won against the steam hammer. If you need a human being in the loop to decide on the allocation of AI tokens, then all those predictions of mass redundancy are gone.
But what if the payoff takes longer than consensus assumes? That question is particularly pressing given that token prices continue to decline and Chinese models are gaining ground, both in their share of the world's most-used models and in token usage, where they now lead their US counterparts among the top 20 models,
Slok provides two charts, the first tracks market share by country of origin monthly since January 2025 among the top 50 models. It is bad for the US, showing a steady erosion of US market share until, eyeballing it, as of May 2026 it is roughly a 60/40 US/Chinese split.
Clearly, the market is voting with its feet that the prices charged by US AI companies are unsustainable.
The second compares monthly token use among the top 20 models by country of origin between May and June this year. It is much worse for the US.In May the Chinese had 80% of the US usage. In June they had 185% of the US usage. If this rate of market erosion were to continue for a few months it would be impossible for investors to continue to imagine the golden future awaiting OpenAI, Anthropic, xAI, Meta, Oracle and the neoclouds.
Slok's analysis of the fallout of the bubble deflating is worth reading. This is an issue I plan to return to in a future post.
Chinese AI providers are making more money; Zhipu’s revenues went up 60 times in Q1 compared to last year; Alibaba is the company behind Qwen, and their revenues are up 15 times just since the beginning of this year. The lion’s share of the profits, though, are still being realized by the hardware side; the chipmakers, at least for now. The model companies’ revenues are rising, and steeply. But so is their cost of compute.
That dynamic is causing some Chinese labs to raise prices; the cost to use Tencent Cloud increased by over five times in March, and Alibaba, Zhipu, and ByteDance also hiked prices.
DeepSeek, however, went the other way. Their latest version is priced at just a fourth of their introductory product, and they rolled out dynamic pricing that is aimed at corporate users, who use their models during the workday.
But even with these price increases, Chinese models just cost far less than those on offer from Silicon Valley, and explains those big jumps in the exports, we can say, of Chinese AI tokens. The LLM’s out of China are “90% as good at 10% of the cost”, and American firms are buying more tokens from Chinese companies. In early 2025, token demand from Chinese LLM’s was about zero. Even the release of DeepSeek didn’t move the needle much, but by the end of the year the secret was out, and it’s been a choppy but steady ride up to 46%, today.
Why would you invest in a company whose competitors were “90% as good at 10% of the cost” and which was rapidly losing market share?
If the margins, and thus the rational valuations, of AI companies are unsustainable, how long can the market remain irrational? In The Second Derivative: Why No One Understands the AI Boom Groundbreaker starts by examining the 2008 crash:
The implicit underwriting assumption, shared by originator and borrower alike, was that the loan would never reach its reset: rising home values would manufacture equity, the borrower would refinance into a fresh teaser and the clock would start again. The structure was a treadmill, and the treadmill was powered by appreciation. It worked spectacularly while it worked. Nearly four in five subprime hybrid ARMs originated in 2003 had been refinanced away by the end of 2006.
Now watch the timing. National home-price appreciation did not crash in 2006. It decelerated. The year-over-year rate of gain, which had run in the mid-to-high teens through 2004 and into early 2005, began bleeding off - still positive, still printing green, but slowing. Prices were higher than they had ever been. And yet, with prices at their peak and still rising, subprime delinquencies inflected upward.
...
The deceleration was endogenous to the structure; the structure required ever-accelerating prices to keep refinancing its way out of its own reset schedule, and no series accelerates forever.
The second derivative was always going to roll over. When it did, the first derivative followed it down through zero, negative equity spread from the margin inward, and the defaults the market insisted were caused by “falling prices” had in fact begun a year earlier, when prices were still rising but had stopped rising faster.
The market is pricing AI as a technology cycle when its actual anatomy is that of a credit-driven real estate cycle - which is precisely why the 2008 mechanics apply - and the two break for entirely different reasons.
...
Walk down the AI build-out and every feature is a property development in disguise: a data center on entitled land, financed with debt against the structure and leased to tenants on take-or-pay terms. This is not a software business that happens to own servers. It is a real estate business that happens to compute.
When a hyperscaler or a neocloud reports a record capex figure, the financial press reads it as confidence, as proof of demand. Read it instead as origination volume. Each gigawatt of committed build is a loan extended to whichever tenant has signed the take-or-pay beneath it, and the credit quality of that loan is precisely the credit quality of the tenant. The market is celebrating loan growth and calling it revenue growth.
Loans in this market take the form of Remaining Performance Obligations (RPOs). The biggest borrower in this market is OpenAI:
OpenAI has committed to pay for compute on a scale without precedent in corporate history: multi-year, take-or-pay capacity contracts whose aggregate obligations run to the hundreds of billions of dollars. Against them sits an operating business that does not yet earn a profit - revenue real, large, and growing quickly, but short of covering the company’s own cash burn and nowhere near covering the contracted payments.
Those payments are therefore not serviced out of earnings. They are serviced out of financing, and financing, for a borrower in this position, is available on a single condition: that each new round price above the last.
Measured as a level, OpenAI’s valuation is the most remarkable appreciation in the history of private markets - roughly $86 billion in early 2024, then about $157 billion, $300 billion, $500 billion, and approximately $852 billion by the spring of 2026. Measured as a rate of change, the same series inverts: the round-over-round step-up ran 1.83×, 1.91×, 1.67×, 1.70×, and falls to roughly 1.23× implied by the reported public-offering target. Private marks are inherently lumpy - negotiated, episodic, set by a handful of insiders - so no single step is decisive. But the trend is unmistakable: it bends down, and it bends hardest at the one mark set by the deepest, most unforgiving pool of capital - the public market. The implied IPO step-up is both the lowest in the sequence and the hardest to negotiate, and it is the one the structure must actually clear. This arithmetic is also the most probable explanation for OpenAI’s recent IPO delay.
In the same way that the private stocks generating unrealized gains are not liquid assets, neither are the Remaining Perfoemance Obligations:
An RPO is not a liquid asset; it is a forward contractual commitment - a promise of future payment in exchange for future compute. And a multi-year commitment is worth exactly the creditworthiness of the entity on the other end of it. When that entity is investment-grade and cash-generative, the backlog is what it claims to be: high-quality visibility, merely deferred. When that entity is a pre-profit company that loses tens of billions a year and can pay only by continuously refinancing its own equity valuation, the backlog is something else entirely. It is a subprime commitment, used to justify massive, un-depreciated capital expenditure, reported to shareholders as structural strength.
Now price the credit quality of that book. Of roughly $2.1 trillion in aggregate contracted backlog across the four big platforms, about half - on the order of $1.05 trillion - is owed by OpenAI and Anthropic. Microsoft’s book is about 49% these two names; Oracle’s is 54%, with roughly $300 billion owed by OpenAI alone; Google’s is 43%; Amazon’s is 51%.
The hyperscaler has, in economic substance, extended a concentrated, unsecured loan to cash-burning tenants. The RPO that Wall Street values as forward revenue is, in reality, a credit exposure to borrowers with no operating income.
Major financial institutions are similarly concerned. In their 2026 Annual Economic Report, the Bank for International Settlements writes:
In the near term, the ongoing AI investment boom raises questions about the sustainability of the current economic expansion. The five largest hyperscalers are set to spend over a trillion US dollars on AI-related capital expenditure from 2025 through 2026. These commitments are outpacing earnings and the free cash flow of these firms, leading some to issue debt to raise additional financing (Graph 11.A). This investment race may be partly driven by the perception that only a small number of players with superior technology will ultimately dominate the market shares. The intense competition raises the risk of firms over-committing resources to investment projects with still uncertain returns, leaving all firms vulnerable to disappointments in AI payoffs. Model analysis based on such contest motives highlights the downside risk of current AI exuberance. As competitive pressure drives capex higher, the net economic surplus – the total payoff less investment costs – declines for the sector as a whole and could turn negative in adverse scenarios (Graph 11.B). Disappointment in returns could trigger a sudden pullback in financing and turn the capex boom into a protracted investment bust, with potential knock-on effects on financial conditions (see below).
Another risk is that the AI boom runs into a supply side roadblock. The AI build-out has recently been facing growing bottlenecks in electricity, advanced semiconductors and grid equipment. Fast-growing demand for computing power is already pressuring electricity prices and input costs, with potential spillovers to inflation. Looking ahead, these temporary shortages may also amplify over-investment, as firms attempt to lock in future capacity through long-dated contracts that further expose them to any disappointments in demand.
Historical episodes of investment booms offer instructive parallels (Graph 11.C). The canal mania of the 1830s, the British railway mania in the 1840s, the electrification exuberance of the late 1920s (roaring 20s) and the dotcom boom of the late 90s all shared one common trait: a genuine technological breakthrough that attracted capital in excess of what commercial returns could ultimately justify. These episodes ended with an eventual reversal in investment, inducing economy-wide recessions. The scale and pace of the current AI investment boom accompanied by expectations of large productivity payoffs bear resemblance to these precedents, highlighting potential downside risks in the near term.
A draft report inside the Treasury Department is set to warn of the risks posed by the artificial intelligence market, likening key aspects of it to the dotcom bubble that upended the U.S. economy when it burst in the early 2000s.
The document, the existence and contents of which have not been previously reported but was obtained by NOTUS, is a significant departure from the Trump administration’s public tone, which has focused on encouraging unrelenting investment to unlock exponential growth.
Career Treasury analysts found that AI firms are more deeply entrenched in the U.S. economy than their dotcom predecessors and pose significant risk to the entire system if financial conditions change, productivity goals are missed or various choke points stymie growth.
Vanderbilt University's Asad Ramzanali has a detailed look at the range of impacts from the burst bubble, and an set of optimistic suggestions for policy responses in After the AI Crash. He frames the problem thus:
Companies are investing trillions of dollars based on tens of billions of dollars in revenues. Analysts at J.P. Morgan anticipate $5 trillion of AI infrastructure investment in the next five years. They estimate that the industry will need to generate annual revenues of $650 billion to justify this level of investment, while consultants at Bain & Co. estimate $2 trillion in needed annual revenues. Yet, OpenAI and Anthropic earned $13 billion and $4 billion, respectively, in 2025 revenues. OpenAI’s own financial expectations suggest negative cash flow until 2030, and Anthropic expects a small profit no sooner than 2029. Alphabet, Meta, Amazon, and others may experience increased marginal revenue from integrating AI into existing products, but that is far from certain.
Note that over the next 5 years around $3T (60% of $5T) of the investment in AI infrastructure goes in to buying the hardware which should be fully depreciated over much less than 5 years. So the gap between the investments and the revenue is much bigger than it appears.
Blackstone's QTS said on Thursday it had terminated its planned Digital Gateway data center project in Virginia and withdrawn the associated filings after years of planning and regulatory review.
The data center operator has faced years of local opposition and litigation over the project, despite it being approved by the Prince William Board of County Supervisors.
Among the hyperscalers, Oracle is the most exposed because, as Ed Zitron noted:
And Oracle ... is a company that, even before the AI bubble, was massively indebted. It just so happens that, as a result of its tryst with OpenAI, Larry Ellison saw fit to twist the debt knob to eleven.
Six firms alone — including Oracle, Microsoft and Meta — have committed $850 billion for data centers leases that haven’t begun yet. Oracle holds the largest share of these commitments owing to its $300 billion Stargate contract with OpenAI.
When Oracle mentions the risk of nonpayment, the unnamed elephant in the room is OpenAI. As part of the Stargate deal with the AI company, Oracle is developing massive data centers across the country to provide cloud computing power. For this plan to work, OpenAI needs to pay its Oracle Cloud Infrastructure bills.
“Some of our customers may be highly leveraged and subject to their own operating and regulatory risks and, even if our credit review and analysis mechanisms work properly, we may experience risks of non-payment and non-performance in our dealings with such parties,” Oracle said in the filing.
SoftBank Group Corp.’s talks with potential creditors to raise at least $6 billion from a margin loan backed by its OpenAI stake have stalled, people familiar with the matter said, just weeks after the Japanese conglomerate cut its initial target from $10 billion.
SoftBank Group has reopened talks with a consortium of lenders for a $10 billion loan backed by its stake in OpenAI, after earlier attempts to secure a loan stalled over concerns about the difficulty of valuing private companies, two people familiar with the matter said.
To make lenders more comfortable, the Japanese technology investor is offering to guarantee repayment of the loan, giving banks recourse to SoftBank if the OpenAI shares pledged as collateral lose value, the people said.
This all seems to indicate that potential lenders, such as banks, are highly skeptical of the value of OpenAI stock.
Despite Grok being so bad that employees use Claude, Musk is touting SpaceX as an AI company. This resulted in the most overvalued IPO in history, which failed to raise enough money to avoid the immediate need to borrow $25B. Nir Kaissar's SpaceX Is Junk. That’s What the Bond Market Says reports on the bond market's reaction:
Ratings companies and the bond market have very different views about how things are going. SpaceX’s bonds have an average rating of BBB across the three majors, Moody’s Ratings, S&P Global Ratings and Fitch Ratings, according to credit scores compiled by Bloomberg. In the alphabet soup of bond ratings, it’s the lowest grade still considered quality before falling into junk territory.
The bond market has other ideas. There, quality is judged by a bond’s credit spread or the additional yield it offers above Treasuries with similar maturity. The wider the spread, the lower the quality. Corporate bonds with a BBB rating are trading at an average credit spread of 0.92 percentage point. SpaceX’s bonds, by contrast, trade at a significantly greater average spread of 1.62 percentage points across maturities, higher than BB rated junk bonds’ average spread of 1.55 percentage points.
The bond market seems to agree with Softbank's lenders about the AI bubble.
OpenAI and Anthropic are competing for the next trillion-dollar IPO. Both would need to distract investors from their massive losses by focusing on growth. Recently, Anthorpic has been growing faster than OpenAI, so Keach Hagey and Berber Jin report that OpenAI Considers Drastic Price Cuts, Anticipating War for Users With Anthropic:
OpenAI is considering drastically lowering the prices it charges users as it seeks to win customers from its rival Anthropic.
The company is weighing significant cuts to what it charges for tokens, the unit of measurement artificial-intelligence firms use to bill for their products, according to people familiar with the matter. The move would be in anticipation of similar cuts the company expects at Anthropic, the people said.
Meta will also introduce a new Meta Model API system, which will be used to collect fees from developers. Its API pricing is roughly 25% of the cost advertised by other top models from OpenAI and Anthropic PBC. Developers will be able to use Meta’s model for free, but only up to a point; they’ll be required to pay for access after reaching a certain token threshold, Zuckerberg said.
“The pricing from some of the other labs is very extreme and has very high margins,” Zuckerberg said, underscoring that his strategy is to get Meta’s technology in front of as many people as possible. “We think that there’s a real ability to be able to offer frontier or very high-level intelligence at a much more affordable cost.”
Aggressive pricing means Meta Model API will still be 50% more expensive than DeepSeek but not 50% better. Planning to reduce current income in the lead-up to an IPO is an unusual move, but it is a response to AI's Affordability Crisis.
Factory electricity bills are generally rising faster than those for other business customers or residential customers, according to a Reuters analysis. It highlighted the example of the Belden Brick Company, a 141-year-old brick manufacturer in Ohio, whose electricity bills have soared from $1,600 to $12,000 per month due to a higher monthly capacity charge in the 13-state region served by the grid operator PJM Interconnection.
...
The Ohio-based steelmaker Metallus described its electricity costs as having jumped by 70 percent since 2024, leading the company to pay an extra $15 million in energy costs annually.
The higher electricity costs for manufacturers coincide with the attraction of large AI data center projects with substantial electricity needs to many states in PJM territory. That data center growth has driven up PJM’s capacity prices—paid to power generators according to supply-and-demand forecasts—from $28.92 per megawatt-day in 2024 to $329.17 per megawatt-day in 2026, according to Reuters’ reporting.
Source
As I see it, there are three separate markets for LLMs. First, there is an embedded market that runs on low-cost, low-power hardware and open-weights models. Small AI Models Gain Traction Around the World by David Berreby provides examples:
For example, a drone-based system developed by Bala Murugan and colleagues at the Vellore Institute of Technology, in India, takes photos of cashew plants and quickly identifies those with splotches that indicate disease. All the processing takes place on the drone itself, so there’s no need for a computer on-site, nor for a connection to a central server.
It isn't just that small hardware, such as the Raspberry Pi 5, is getting more powerful but also:
the shrinking footprint of language models. Both Google DeepMind’s Gemma 4 (released in April) and Alibaba’s Qwen 3.5 are “fantastic” for small AI, Rovai says. Both models are “open weight,” meaning users can adjust the connections between parameters to suit their needs. This makes it easy, for example, “to take a lot of data from, say, the milk industry and retrain the model specifically on that,” Rovai says.
The hyperscalers and AI platforms like OpenAI and Anthropic will garner no income from this market, because:
“I think the future of AI is not like one giant model, at a center. I think it’s millions of small, precise models deployed at the edge, each one solving like a specific problem, a specific context,” Alonge says. This is partly because much of humanity—including people in parts of rich countries as well as the developing world—lives without access to cutting-edge frontier models. But, he says, it’s also because those models are not sustainable.
“If someone is not subsidizing it, most people will not be able to afford those models. So those of us who are said to be small-AI developers are the ones who will have to build for the majority of the world,” Alonge says.
At Computex 2026 in Taipei on June 1st, CEO Jensen Huang announced the RTX Spark superchip — a single piece of silicon that combines a 20-core Arm CPU, a Blackwell GPU with 6,144 CUDA cores, and 128 gigabytes of unified memory, connected by NVIDIA’s NVLink chip-to-chip interconnect. The whole package delivers up to one petaflop of AI compute in a laptop form factor.
The number that matters: RTX Spark can run a 120-billion-parameter language model entirely locally, with a context window of one million tokens, without a single byte leaving your machine.
To put that in perspective: GPT-3 had 175 billion parameters and required clusters of A100 GPUs to run. The model that stunned the world when it launched in 2020 is now approximately the size of what fits in a consumer laptop chip announced this week. The capability that required a data centre in 2020 is coming to a device you carry in a bag in 2026.
In 2025, slightly more than a third of all smartphones shipped worldwide were capable of running generative AI, and that figure will reach 45 percent by the end of this year, according to the technology research firm Counterpoint. By the end of next year, slightly more than half of all smartphones will be able to run a small AI model.
If Nvidia can already put a 120B-parameter in a laptop, it will only be a few years until phones can run a GPT-3-class model, good enough for almost all consumer needs. Apple and Google own that channel. Owning the channel is better than owning the technology. They will dominate consumer AI, and the other players will garner no revenue from this market.
Third, there is an enterprise market; all that is left to generate the revenue to service the debts fuelling the AI bubble, lets say $2T by 2030. There are a number of problems that make this unlikely.
First, there are very few documented cases of LLM deployment that resulted in enough productivity improvement to cover its unsubsidized costs.
Second, the Trump administration just demonstrated that deploying mission critical systems on AI platforms such as OpenAI or Anthropic means your company can be disabled at 90 minutes notice with no recourse.
Third, this means that companies will have to run mission-critical LLMs on open-weight models on in-house hardware if they are not to be vulnerable to the whims of the US president.
Fourth, systems such as RTX Spark show that good enough in-house hardware is likely to become relatively cheap compared to the unsubsidized cost of the AI platforms. In-house systems need much less over-provisioning for demand spikes, and because they aren't shared they don't need to be as fast.
Fifth, companies need to balance the productivity benefits (if any) of mission-critical LLMs against the productivity costs they bring. These include a vastly greater attack surface, technical debt from reduced developer understanding of the software, and so on.
Thus it seems likely that the hyperscalers and AI plaforms will generate far less revenue than they expect, because they will be restricted to non-mission-critical applications with lower productivity gains, and thus lower pricing power. They will thus be unable to cover the debts they are incurring to build massive data centers predicated on centralized systems dominating (an inflated estimate of) the entire enterprise market.
Karp had been softly floating his critique for some time, but the CNBC event looked like a proper coming out. Just one day earlier Palantir had published a kind of manifesto devoted to what it described as the all-important principle of “A.I. sovereignty.” The central argument: Companies should seek to build their own A.I. tools, not just customize those on offer from the frontier labs. This might mean relying on open-source L.L.M.s rather than the proprietary ones on which the A.I. boom has mostly been built in America, but it would amount to a liberating declaration of independence from Big A.I., which in Karp’s estimation was sucking up much more value than it was generating.
Locking in a price with a multi-year neocloud contract insures a buyer against compute getting more expensive but not against it getting much cheaper. So as businesses around the world spend ever more on compute, they want to be able to hedge against price volatility just as they insure against changes in energy tariffs, interest rates or foreign-exchange movements—ideally in deep and liquid derivatives markets.
Two startups want to help companies do this, by turning nascent indices tracking compute costs into a futures market. Silicon Data, founded in 2024 and backed by DRW, a trading firm, has paired up with CME Group, which operates large derivatives exchanges. Ornn, created by recent graduates of the Massachusetts Institute of Technology and run from a flat rather than an office just a few months ago, has paired up with Intercontinental Exchange, the parent company of the New York Stock Exchange, to do the same. Both aim to launch compute futures later this year, to be traded on their partner exchanges.
One characteristic of bubbles is overbuilding the infrastructure. Signs of overbuilding of data centers include that both SpaceX and Meta are now in the business of rentling their GPUs to the competition.
As morale is hitting rock-bottom, his company is heavily relying on its competitors' AI models to build out its own in-house tools. And despite the many billions of dollars the company has spent in its flailing efforts to keep up in the AI race, even Zuckerberg himself is now acknowledging that progress is nowhere near where he wanted it to be.
As Reuters reports, Zuckerberg admitted during a town hall last week that AI agents in particular aren't progressing as fast as he anticipated, a devastating revelation following enormous layoffs that wiped out thousands of roles at the company.
The "trajectory of the agentic development over at least the last four months hasn't really accelerated in the way that we expected," he said according to a recording obtained by Reuters.
Despite a derisory ~3% share of the enterprise LLM market, SpaceX's IPO was marketed as an enterprise AI company. The IPO has to rate as the most manipulated of all time featuring, in addition to ludicrous financial projections, a tiny float, bent rules for index inclusion, and massive conflicts of interest at the banks and the analysts. Immediately afterwards, SpaceX issued $25B in bondsa. A month later we can see how the markets respond to the first trillion-dollar "AI company".
We’re old enough to remember when the market cap of the lossmaking telecom SpaceX was bigger than Amazon’s. Heck, for a few precious moments it was bigger than Microsoft’s. Maybe one day it will be again, but for now the stock is down 38 per cent from its peak post-IPO valuation.
Punters lucky enough to have been awarded a stock allocation at the outset are still sitting on a tasty [checks notes] 0.8 per cent paper profit at pixel time.
And this is before the lockups start expiring and the initial tiny float greatly expands.
SpaceX bod spread
Second, more interesting as being much less subject to manipulation, are the bonds. They are what Toby Nangle focused on in SpaceX bond yields rocket towards junk:
the full $25bn of benchmark bonds — issued across the curve — had a rocky first couple of days of trading. Checking back today, it turns out that the inauspicious beginning was just a prelude to the train wreck that has since unfolded.
...
If you’d been allocated $100mn of the SpaceX 2056 bonds, you’ve turned $100mn into $90.7mn in less than a month. Sure, long-dated US Treasury bonds have fallen in value, and this general sell-off at the long end has done some of the work. But the spread on SpaceX 2056 — the additional yield you’re paid to compensate you for the risk that you don’t get repaid (among other things) has now widened from the initial +175bps to a whopping +231bps doing more than two-thirds of the work.
Top 10 worst BBB
The same conflicts of interest that pumped the stock caused rating agencies to give SpaceX bonds a BBB rating, one notch above junk and crucially the lowest that many major institutions are allowed to hold. But Nangle notes that:
Looking only at the nine days since the bonds were included in ICE BofA indices at the end of June, this spread-widening has made SpaceX 2056 the single worst-performing US dollar triple-B benchmark bond:
Note that the top two worst bonds are SpaceX, but the rest of the top 10 are all Oracle, another financial disater area.
when we overlay the average spread for double-B US dollar corporate bonds across different maturities (the pink line), it looks a lot like the type of risk that the market has assigned to both SpaceX and Oracle bonds is junk risk.
SpaceX, OpenAI, Anthropic, Meta all need to raise vast amounts of debt to fund their plans for AI. With SpaceX's bonds trading as BB despite a BBB rating, this is going to be hard.
Founded in early 2025 by former OpenAI CTO Mira Murati, Thinking Machines' first model is a big one. Weighing in at 975 billion parameters, the model requires more than two terabytes of GPU memory — a quantity present in around eight of Nvidia's B300 accelerators, or sixteen H200s — to run at its native 16-bit precision. If that's asking too much of your hardware, Thinking Machines has also released a NVFP4 quantized version of the model capable of running on half the GPUs.
This makes it the largest American open weights model to date, and comparable to Chinese models like DeepSeek V4, GLM 5.2, and Kimi K2.6 in terms of size and capabilities. Take these claims with a grain of salt — gaming AI benchmarks isn't exactly difficult – but Thinking Machines says Inkling is competitive with these models in a variety of workloads, although its benchmark charts also show it trailing proprietary models like Anthropic's Claude and OpenAI's GPT.
If you don't like Chinese open-weight models, the US ones are getting better:
The model developer claims to have tuned the model to use these thinking tokens more efficiently and that Inkling therefore matches Nvidia's Nemotron 3 Ultra, up to now the largest and most capable American open weights model out there at 550 billion parameters, on Terminal Bench 2.1 using roughly a third the tokens.
The AI revolution is supposed to be in its early innings. Demand for GPU computing is exploding. Microsoft has relied on CoreWeave for massive amounts of AI compute. OpenAI has signed tens of billions of dollars in long-term contracts. NVIDIA isn’t just supplying the chips, it has also invested in CoreWeave and entered into agreements that support parts of its financing and capacity.
So why are CoreWeave insiders continuing to sell stock? Management says many of the sales were made under pre-arranged Rule 10b5-1 trading plans. But those plans explain how the shares were sold—not necessarily why executives continue converting stock into cash while investors are being told AI infrastructure demand has never been stronger.
New York has become the first U.S. state to stop construction of large new data centers, imposing a one-year moratorium due to growing concerns over power costs, water supplies and the burden on local communities.
"As data center development threatens to hike up utility bills, deplete our natural resources, and create uncertainty for New Yorkers, it's my responsibility to take action and lead," said New York Governor Kathy Hochul.
She added that she would also pursue legislation to repeal sales tax exemptions for large data centers.
Phurichai Rungcharoenkitkul continues the Bank for International Settlements's pessimism in The AI Investment Race:
The AI build-out ranks among the largest technology-driven investment booms in US history. Its scale, reliance on debt and circular equity ties raise questions about the boom’s sustainability and financial stability. We study a dynamic contest in which firms competing for a few dominant positions over-commit resources. The over-investment leaves the sector exposed to revenue disappointment that could turn boom into bust. The larger the boom, the deeper the eventual bust. The race to commit early through debt and circular financing also makes a bust more likely. Calibrated to balance sheet and deal data, the model points to over-investment of around 1.5 times the efficient level, rising to around three times where demand is less elastic. A network analysis shows that stress in one firm could cascade to others through chains of financial exposures.
As recently as this week, one executive at Anthropic PBC, who spoke on condition of anonymity, mused that the Claude maker’s technology was roughly six to 12 months ahead of Chinese rivals.
On Friday, Moonshot AI Inc. upended those assumptions. The Chinese AI lab released Kimi K3, a more advanced open-weight model that it said outperforms all rivals except for Anthropic’s Claude Fable 5 and OpenAI’s GPT-5.6 on overall capability. The implication is that Moonshot, and by extension China, could be closing the gap faster than expected.
...
Moonshot’s release could also undercut OpenAI, Anthropic and others on price at a moment when they’re confronting customers who are becoming more conscious of their surging AI spending. Some developers have turned to so-called model routing services that can seamlessly direct users to cheaper options, including from China, for specific tasks to maximize cost efficiency.
Kimi K3 is an open-weight model.
If a 120B parameter model is too small for you the RTX Spark won't cut it. But Michal Malewicz explains in NVIDIA just killed big AI and… You’re the winner? that one of Nvidia's OEMs will sell you a $94K box that will run a 1T model, the Exxact Valence Nvidia DGX Station.
"[O]ur bottom line is that AI adoption appears gradual and steady rather than rapid and transformative, with most households and businesses still reporting limited exposure to the technology. At the same time, evidence of a structural pickup in productivity growth remains surprisingly fragile: Aggregate productivity growth has improved during the post-pandemic period, but there is a strong rationale to attribute much of this improvement to cyclical variation in utilization rather than a sustained acceleration in productive capacity. We also find little compelling evidence in the available industry-level data that industries adopting AI more rapidly are already experiencing stronger productivity growth."
"productivity trends across all three levels have been relatively consistent over time, suggestive of micro-level productivity gains not adding up in aggregate"
Another instance of Betteridge's Law is Will OpenAI ever be profitable? from Leap Finance Academy. It is a very long and detailed examination of all the things that would have to happen for OpenAI to turn a profit by 2030. In summary, with my comments, the list is:
Reduce cost of inference which isn't going to happen enough because each successive frontier model uses more tokens to generate the same output.
Increase pricing which isn't going to happen because they already need to reduce pricing to stem the loss of customers to Anthropic, let alone to the open-weight models.
Increasing sources of revenue that are not inference dependent in other words the fantasy of $100B/year in advertising income, which isn't going to happen.
Securing and retaining human talent which requires an IPO to make their stock options worth something, which isn't going to happen it time to prevent bankruptcy.
Securing the funding they need to build Stargate and for working capital which isn't going to happen because even an IPO wouldn't come close to generating enough and the lenders are saying "enough".
Even though I'm skeptical of the wishful thinking of the conclusion, this is an impressive piece of work and well worth reading.
China’s open-source approach has spawned a crowded field of innovative and intensely competitive start-ups all offering systems at low cost.
So attempts to increase revenue by charging more for access to certain models can scare away customers. Price-conscious Chinese consumers — businesses and individuals alike — are quick to hop across platforms in search of inexpensive A.I. tools.
While China’s A.I. companies are searching for sustainable business models, spending is high and revenue low, said Richard Lin, a vice president at the Silicon Valley company Datastrato.
“In two or three years, we will still be trying to figure out how large models can earn money,” he said.
Offering low prices has helped the Chinese firms gain users, including in Silicon Valley, where many companies depend on the more affordable systems. Yet Chinese companies have struggled to translate huge numbers of global users into profits.
Bain & Company management consultant Jue Wang said many of the big businesses her firm advises have been taking a closer look at returns on their AI investments.
“The token cost for them has been doubling, almost every other month,” she said. “Let’s say $200 per developer per month. Multiply that by 20,000 developers, which is often what we’re dealing with at these companies, and that quickly gets you to a number that is not a line item that any general manager has planned for.”
Sometimes that just means not using the AI equivalent of a sledgehammer to crack a nut.
A huge wave of supply of debt intended to fund the data center build-out for AI means that the price of this debt has dropped, and thus that the interest rate buyers are demanding goes up. This reduces the net pressent value of the (hypothetical) future cash flows the data centers will generate, and thus their ability to pay the interest and principal. This increases the interest rate the buyers of future debt will demand, causing a feedback loop.
Both Ed and Torsten point to the increasing spreads above Treasuries that the AI debt is trading at, and the increased cost of insuring against these companies defaulting.
In a blink, [Leopold Aschenbrenner's] wildly successful hedge fund, Situational Awareness, was forced to sell billions of dollars of technology investments that had rapidly lost value, as nervous banks began to demand more and more collateral for his trades. Then came billionaire Ken Griffin.
In less than 24 hours — which included a conversation between Griffin and Aschenbrenner — Griffin’s Citadel hedge fund reached out to Situational Awareness and snapped up the investments at a discount, according to a person familiar with the matter who asked not to be identified citing private information.
It was a startling reversal for Aschenbrenner, a former researcher at OpenAI who — before starting his hedge fund roughly two years ago — had no previous investment experience. His fledging firm has watched its assets plunge from $45 billion at the start of July to about $10 billion.
Aschenbrenner was all-in on AI.
On the Pof. G. Markets podcast legendary short-seller Jim Chanos focused on the accounting inequality inherent in the AI (and earlier dot-com) bubble. The spending of the hyperscalers and neoclouds is investment, expensed over say 5 years. But for the chip makers, construction companies and so on that spending is this year's revenue. So, overall, it looks like the ROI is better than it really i
s.
US GAAP was built on the assumption that neither route would materially move profits, because which company would ever choose to grant ‘enough’ RSUs to employees that it could move the share price right?
Meta has granted ~US$70bn in stocks. And Microsoft ~US$42bn. The sheer size of these numbers IS enough to move the needle materially. RSUs, granted in large quantities every year, compound into visible profit erosion and free cash flow consumed by “obligatory” share buybacks, as employees must receive their promised shares annually regardless of market conditions.
As of FY25, total unvested RSU’s book value was ~US$60bn and corresponding market value was ~US$80bn. That $20bn gap is money Meta will have to find, one way or another, the moment these shares vest, either by diluting shareholders with fresh stock, or by spending real cash buying shares back to hand over instead.
...
Using the extensive information provided by Meta’s financial notes on the RSU programme, we can back-solve the profit erosion that is being kept ‘off-the-Income-Statement’:
Since December 2024, as the share price of Meta increased and the AI talent war started, ~10 points of EBIT margin per year have been given away as additional, uncaptured employee compensation.
Ed Zitron is back with another scathing analysis in The AI Demand Bubble. As usual, it is long and detailed. He concludes:
To put things really simply, Anthropic and OpenAI are a way that hyperscalers can feed their revenue to themselves by spending money on capex, backstopping compute contracts, or doing direct equity investments.
Their continued existence allows the AI bubble to continue inflating, but this can only continue as long as venture capital and hyperscalers are capable or willing to invest. There is simply not the demand — not from open source, not from other AI labs, not from self-hosting, not from anywhere — to justify the capex or the massive data center buildout.
And for those arguing that there would be a dot-com bubble recovery story, I must be clear that if there isn’t demand today, it won’t magically appear tomorrow. AI GPUs will cost just as much to run in five years as they do today, as will unfinished data centers cost just as much to finish, as will electricity remain expensive, and all this will be happening after it’s easy to raise venture capital to actually buy the compute.
In a recent survey of 300 business executives, 68% reported overspending their AI budget over the past year. It’s little wonder so many companies are now cracking down. Uber Technologies Inc. recently capped employees’ AI spending at $1,500 a month. Tesla Inc. reportedly set a limit of $200 a week. The swift reversal has left some workers complaining of whiplash, with Reddit forums filled with stories of engineers suddenly stymied by usage caps.
Bloomberg in partnership with researchers at Vals AI, an independent AI evaluation and benchmarking platform, tested seven models from frontier Chinese and US companies to see how they performed in a real-world task. They were asked to create a fictional coffee e-commerce site called Brewberg using the same prompts. Most of the models scored 100% functional accuracy despite occasional design misses, but with very different price tags. The experiment employed the top performing models in July from Anthropic and all the Chinese firms, as well as more affordable models from OpenAI and Google.
They all did reasonably well, but the two best were Claude Fable 5 at $48.99 and Kimi K3 at $11.99. Chinese models charging much less for almost the same performance are grabbing market share:
the use of Chinese models overtook US platforms globally for the first time in June, and accounted for more than 60% of market share last month, on OpenRouter, a tech platform that offers software developers access to hundreds of AI models. It is a widely watched gauge of model usage despite tracking just a fraction of global AI consumption. The US, parts of Europe and Asia now favor Chinese labs, according to the same data.
On Hugging Face, Chinese AI models account for 41.4% of generative model downloads among developers, 5 percentage points higher than US models.
In OpenAI’s unraveling has begun Gary Marcus notes that leaked Q2 numbers show OpenAI's quarterly revenue grew from $5.7B to $6.7B but its quarterly losses grew from $9.3B to $12.3B.
Perplexity is launching Portable Computer today, a version of its agentic "Computer" platform that runs entirely on hardware users already own — starting with Nvidia's DGX Spark desktop supercomputer and Linux machines equipped with Nvidia RTX GPUs.
The launch, developed in close partnership with Nvidia, is one of the most aggressive attempts yet to move serious AI agent workloads off the cloud and onto local devices. The model, the user's files, and the work itself can all stay on the machine. Work completed locally consumes no billing credits, and the company says every task starts on the device by default — with the system asking permission before sending any individual step to a more powerful frontier model in the cloud.
Apple announced new iterations of both desktops, along with two new chips: the M6, the first 2nm chip in Apple’s M-series lineup for Macs, and the M5 Ultra, now the most powerful chip in the lineup for most things—especially AI workloads.
There aren’t any major new features for either machine. This is just a specs bump. But based on how Apple is presenting these refreshes, they’re leaning hard into those use cases, which weren’t even a thought when earlier iterations were first engineered.
The devices’ popularity for production inference took off after macOS 26.2 shipped last December. According to Apple’s release notes, 26.2 enabled “low-latency communication between Thunderbolt 5 hosts for use cases including distributed AI inference using MLX.” Thunderbolt 5 is a very fast wired data connection, and MLX is an open source array framework designed to help machine learning workflows take full advantage of the M-series chips’ unified memory.
Since then, both hobbyists and professional developers and researchers have been essentially daisy-chaining Mac minis or Mac Studios to run inference on local large language models that are much bigger than anything that could run a single mass-market device—providing an alternative to ultra-beefy specialized hardware featuring specialized Nvidia GPUs.
IBM has rolled out the newest models in its family of open-weight large language models designed to be downloaded and self-hosted. The newly launched Granite 4.2 comes in 3B, 8B, and 30B parameter variants.
Like previous versions, IBM is taking a decoder-only approach here. These new releases offer a 128,000-token context window natively. The 8B and 30B variants (not the 3B one) also go through an agentic reinforcement-learning block; they were trained for expanded capabilities like using the terminal, searching the web, or using external tools. The 3B model supports tools too, but without the same level of specialized training.
Beyond those tweaks, this release is particularly notable because, as IBM itself writes, “Granite 4.2 is the reasoning-focused release of the Granite language-model family.
Morgan Stanley analysts today revised their model for how much power would be needed to switch on the chips it projects will be sold. They now estimate, based on a total data-centre power requirement of 257 GW by 2028, a gap of between 30 and 40 per cent between US capacity and sales projections. Or to put it in layperson’s terms:
New York City demands 5.5-6 GW of baseload power: we are short power by six New York Cities.
And that’s using their mid-estimate for innovative “time-to-power” solutions. Absent such mitigation measures, the estimate rises to 10 NYCs:
The innovative solutions include a massive expansion of Bloom Energy fuel cells, data centers alongside nuclear power plants, and other equally optimistic concepts.
But Morgan Stanley are equally optinmistic about Nvidia:
Fenyman, Nvidia’s next next chip architecture, is scheduled for launch in the second half of 2028. It’ll be another step change, Morgan Stanley says: the same dynamics will increase US power demand year-on-year by a further 40 per cent in 2029, Byrd and team estimate. They’re then looking at a cumulative power shortfall of 107GW (or approximately 18 NYCs).
“However,” the team add in bold, “all of these advancements converge on one central issue: if we cannot close the power shortfall, the expected pace of improvement and adoption may not fully materialise in line with market expectations.”
They probably don’t need to spell out what happens to all those big tech off-balance-sheet guarantees in that scenario. Obviously, the faster-rising thing will eventually overwhelm the slower-rising thing,
I think "market expectations" are out over their skis.
I was going to write and publish this post earlier, but the last couple of weeks have been trying to enjoy (the rest of) my vacation and submitting job applications. It’s also been a bit weird. I’ve never worked anywhere else longer than 2 years (mostly because of the jobs being contracts), so it’s sad … Continue reading "Reflection: The end of 8 years at GitLab"
I begin by addressing how we might define public data. It is, in short, data whose value is not determined by the marketplace, by its commodification, or by what Marx would call its “use value.” Public data may function in all of these ways, but its value is defined by something more abstract and less quantifiable. It serves civil society, and without it, an intrinsic public good—the health and safety of a polity, the protection of an environment, the advancement of creative and critical thinking—would be lost.
As stewards of the past and curators of the present who work in service of the future, we in libraries, archives, and museums have long turned our attention to cultural heritage—both in its material and digital forms—for their value as public goods. In the United States, we have, since the first Census in 1790, more or less “Render[ed] therefore unto Caesar” the production and dissemination of information produced by the government for the governed.
Since 1895, civil society in the form of libraries have made this information accessible through the Federal Depository Library Program, which self-describes as the “nation’s link between the American public and its government.” In the digital world, with uneven success, this distribution network has been replicated online through the aggregator Data.gov.
As James Jacobs and Jim Jacobs point out in their recent comprehensive study Preserving Government Information: Past, Present, and Future, “Although there is a general consensus that the government has an obligation to preserve its own information … the laws that affect preservation of born-digital Public Information are outdated and inadequate.” The last year and a half have done more to reveal the vulnerability — the precarity even — of public data as vital infrastructure than any other time in our country’s history.
We are among the collective effort to change the cultural climate around government data so that it is understood as a resource held in the public trust.
In response to this moment, Harvard Law School Library’s Public Data Project began with copying 311,000 datasets from Data.gov between November 2024 and January 2025. We now build cutting-edge tools and interfaces that enable access, discovery, and monitoring of large public datasets. We are among the collective effort to change the cultural climate around government data so that it is understood as a resource held in the public trust. We are grateful for the support of the MacArthur Foundation and the Rockefeller Brothers Fund, without whom we could not do this work. The opinions and views shared here do not necessarily state or reflect those of our contributors.
As part of that collective effort, in August of 2025, we copied 710 TB of public domain data from the Smithsonian Institution — the complete open access portion of the Smithsonian’s collections. We have updated it weekly since then, and now have approximately 828 terabytes (9.1 million files) sourced from more than 20 libraries, museums, and research centers across the Smithsonian. It is this work that I will focus on today.
We have chosen Source Cooperative as the ideal repository for a number of reasons, not all of which I will get into right now. I will just say that because it is built on cloud object storage, the repository supports direct publication of massive datasets, making it easy for us to add, update, and share data. Source also provides a handy user interface for browsing and previewing, along with an API for fast programmatic access.
At present, this repository mirrors the directory structure used by the Smithsonian, but navigating it has proven a challenge. More than 17 million metadata records, constituting over 47 GB, are included, but associating them with the images in this collection, let alone querying them for a particular subject, is difficult. We have therefore generated lightweight Parquet search catalogs for this collection, which can be found under the search directory.
And it’s being used! Early analytics suggest a considerable interest in this largely unmediated data as data. But for us, these 17 million metadata records and their accompanying images are the tip of the iceberg — for two reasons.
First, we position this work among the larger movement to call attention to the importance of federal data as vital public infrastructure, for which allies such as America’s Essential Data advocate by highlighting “how everyday Americans are benefiting from specific federal datasets.” As Fran Berman recently wrote in The Conversation, “society … need[s] to change both its expectations about digital critical infrastructure and its actions — the way the public and private sectors create and regulate these services and systems and the way the public uses them — to better protect the people who depend on them.”
Not only do we position the Public Data Project among such efforts, but — and this gets to the second reason for our work — we position the Smithsonian’s public data as an excellent ambassador for this mission. Want to get people to care about public data? Lure them in with the cool stuff … like this ceremonial Chinese wine vessel from the Zhou period at the National Museum of Asian Art, which can be viewed as a 3D model here.
Zhou period ceremonial wine container (fangyi) from the National Museum of Asian Art. Source: Smithsonian Open Access Archive.
To enhance discovery and monitoring of large government datasets through the building of interfaces and tools, we need collaborators. We could not be happier to have begun, this summer, an intensive collaboration with Kyle Deeds, Ben Lee, and their graduate students Akshay Mehta and Ying-Hsiang Huang.
As a collaborative, we maintain that the heterogeneous and interconnected nature of government data requires new models of search, and this group’s work with vision language and frontier models is leading the way to meet these needs. Together, we are building not just tools, but also a community of experts who will lead the pedagogical and technological interventions necessary to make a real difference in the cultural and political landscape.
As a team of engineers, librarians, and scholars nested in Harvard Law School Library’s Library Innovation Lab, the Public Data Project, in concert with our collaborators, seek — as our mission states — “to grow knowledge and community by bringing library principles to technological frontiers.” In this instance, we see our work with the Smithsonian public domain data as an opportunity to think differently about library, archive, and museum (LAM) catalogs as we have known them.
The word “catalog” is Greek, combining the word kata, which means “completely,” and legein — “to say, count, or gather.” We can understand LAM catalogs as counting or gathering completely, or at least aspiring to gather completely. Objects in collections — paintings, incunabula, newspapers — are the organizing principles by which catalogs are composed. A record is created for each item and then the records are “gathered” or aggregated into a list, an index drawer, an OPAC. And in the world we have known, the institution that holds the object, or its digital instantiation, creates that record and maintains that collective catalog.
Catbriar and huckleberry in the author’s backyard. Photograph by Nick Anderson.
But this story is changing. In our networked and media-saturated world, commentary, both expert and amateur, happens beyond the walls of the brick and mortar institutions that house the objects. I am thinking here of the work of Michelle Caswell and the idea of “community archiving” because, I want to posit, such “bottom up” descriptions offer us a challenge — and an opportunity — to reimagine catalogs.
We want to engage social media as sites of interpretation and cultural mediation.
Object-based records could still rule the day, but the sources for the accompanying metadata might become varied, so that institutionally created metadata sits alongside socially created metadata, what George Oates simply refers to as “social metadata.” We want to engage social media (and other sources created beyond the confines of the institution) as sites of interpretation and cultural mediation.
And in so doing, we want to enhance our understanding of what constitutes a collection, so that we are layering on two large datasets that pertain to the Smithsonian Institution collection and are created by the public.
“Social metadata” becomes part of how we describe and interoperate a collection. We begin with the Smithsonian Transcription Center, which hosts crowdsourced transcriptions of digitized Smithsonian materials. Established in 2013, the center supports volunteer transcription projects across the Smithsonian’s museums, libraries, archives, and research centers. Between July 7 and August 10, 2026, the Public Data Project collected the source files, transcription files, PDFs, metadata, and project manifests associated with 14,167 transcription projects — 854 GB of data in total.
The public cannot rely on industry platforms to see and study social media. The lack of ability to capture social media is one that permeates the digital archiving world.
In addition, we currently have two sources for social media data. First, the Smithsonian Flickr Data Lifeboat, which includes 3,308 photographs; 7.8 GB of data; 5,768 comments; 10,954 tags; and contributions from 4,710 users. We are also an early beta tester for Common Data, a service that will provide an easy way to see and study social media in real time. This platform and data trove are necessary because the public cannot rely on social media industry platforms to help study it.
The lack of ability to capture social media is one that permeates the digital archiving world. If Common Data is transforming that, we believe that this Smithsonian data project is well-positioned to showcase an early adoption of social media archives to enhance government data collections. Common Data is currently developing a system for collecting and preserving social media data from platforms including Bluesky, Telegram, TikTok, and YouTube as well as podcasts, and soon it will include data from Instagram. In the coming months, we will be calling for testers of this new interface and the fine-tuned LLMs and vision-language models underlying it to make connections.
The tech industry likes to talk about how AI enables us to know or capture all or total knowledge, and we might question that assertion as a receding horizon at best, a shibboleth at worst. But, when it comes to catalogs, I’d posit that it behooves us to strive for a deeper record, to “list completely” as the etymology of the word itself demands.
By supporting multi-modal search across different collections, institutions, and types of data — visual, temporal, geospatial, social — we hope to enable search across paradigms as well, different understandings of our past and present, different populations and experiences, and to surface sublimated and hidden connections.
This post was written by members of the DLF Assessment Interest Group’s (AIG) Metadata Working Group (MWG). Learn more about theAssessment Interest Group.
It was authored by: Xiaoli Ma (University of Florida), xiaolima@ufl.edu, Stasha Gardasevic (University of Hawaii at Manoa), gardase@hawaii.edu, Anna Goslen (University of North Carolina at Chapel Hill), goslen@ email.unc.edu and edited by: Anna Goslen
Electronic theses and dissertations (ETDs) often have specialized workflows and metadata management requirements due to factors such as varied ingest methods and inconsistent source metadata. This blog post features descriptions of ETD metadata management from three institutions in the United States, outlining the systems, tools, and workflows used to process and provide access to ETDs at each institution. This post is intended to help other metadata professionals working on ETD metadata management by providing brief descriptions of how the work is handled at other institutions. A list of additional resources on ETD metadata management is provided at the end of the post.
Name: Stasha Gardasevic
Organization: University of Hawaii at Manoa
Once a semester, we get a ProQuest dataset of ETDs to ingest into the Institutional Repository.
Our graduate division uses ProQuest ETD Administrator, where the students log in and upload their submissions. These submissions are then verified/approved by the graduate division.
Three months later, ProQuest sends us a batch of zip files that includes metadata in XML format together with all the files — one zip file per submission. They upload the batch to our file server, with our permission. Our repository admin converts their zipped files and metadata into a format we can upload to our DSpace instance (called ScholarSpace).
There are several things to look out for when importing, as minor corrections might be necessary. For example, student-supplied indicators might be wrong, e.g., D.P.H. (Doctor of Public Health) when it’s actually a Ph.D., or M.A. instead of an M.S. The DSpace collections are determined by those fields, so they need to be corrected; the error will appear when our script can’t find the collection. Submitters can apply an embargo of 6, 12, or 24 months. Those are noted in the metadata and applied when the file is uploaded. Some degrees opt for a permanent suppression, in which case we get only the opening page and the intro.
Submitters can also opt out of “3rd party searches” on the ProQuest form. Since all of our files are crawled by Google, we handle this by putting them behind authentication.
Prior to ingestion, the metadata librarian checks files to make sure the:
Titles are in sentence case.
Personal Names are in preferred format (Last name, First).
Supervisors’ name entries per batch are normalized; supervisor names may occur in strange shapes, and occasionally ETDs authors will use their name for supervisor as well.
Data in other fields appears as expected.
Recently, we have been testing an automated way to make these modifications using Claude Pro. We are working on integrating about 8,000 ETD records that are currently not in our database but are available via ProQuest and were downloaded via TDM Studio (trial). Finally, we are developing an ETDs Trends Exploration dashboard that works on the enriched metadata for previewing trends in topics, supervision, and other stats.
Name: Xiaoli Ma
Organization: University of Florida
At the University of Florida (UF), Electronic Theses and Dissertations (ETDs) are integrated into the Institutional Repository, one collection of the UF Digital Collections. The workflow begins with the Graduate School sending raw XML dumps to Libraries IT, where they are transformed into MARC XML format and uploaded into Alma. Following this, staff in Metadata for Digital Initiatives extract the information from Alma into Excel using the Alma lookup tool, structuring and preparing the data for ingestion into the UF Digital Collections and subsequent distribution to ProQuest. This is the general picture.
Challenges lie in this general picture. First, the locally developed tool that converts Graduate School XML files into MARC XML has problems with field mapping, data formatting, and boilerplate text. Secondly, we receive additional batches of theses, dissertations, project outcomes, and student write-ups outside of the Graduate School pipeline. Arriving throughout the year, these varied submissions require separate preparation workflows for ingestion. Although a single unified metadata standard would be ideal given the similarity of the content, the different intake routes result in varied metadata standards across data batches in practice.
While understanding the benefits of using a single unified metadata standard, we have a hard time moving in that direction.
Our outdated digital platform presents major obstacles to implementing it:
Lack of batch update capabilities:The system lacks an effective, built-in batch update function. Adopting a unified standard would require updating all legacy content, making this an overwhelming task without sufficient support from the system.
Technical and testing hurdles: Even if the new standard were applied solely to incoming content, Libraries’ IT would need to modify code and conduct multiple rounds of testing to ensure the digital platform takes in the XML files.
Insufficient documentation and unresolved errors: This process is further complicated by a lack of documentation regarding the platform’s validation rules, leaving us without clear guidance on what triggers errors. Consequently, many validation errors encountered over the years remain unresolved.
UF plans to transition to a new digital collections platform next year, which we hope will enable us to implement unified metadata standards across all ETD batches.
Name: Anna Goslen
Organization: University of North Carolina at Chapel Hill
Like many institutions, the University Library at UNC Chapel Hill receives a batch of ETDs from ProQuest once a semester which are uploaded to an FTP server. Our institutional repository, the Carolina Digital Repository (CDR), uses the Samvera Hyrax system. When we began using Hyrax in 2019, ProQuest metadata was mapped to our local metadata profile for ETDs and a script was written to transform metadata to the required format. We added additional functionality to our administrative dashboard that allows repository managers to ingest packages with an ‘Ingest from FTP’ button. Part of the metadata transformation process involves mapping student affiliations to our local controlled vocabulary for university departments. Periodically, we are required to adjust the mappings or controlled vocabulary to account for new or renamed departments at the university. After the repository manager ingests new batches of ETDs, the metadata librarian is notified and performs quality assurance to check for any issues with the source metadata or transformation process. The access settings selected by the student are also automatically applied during the ingest process. Our collection of ETDs is dated from approximately 2006 to present; occasionally, pre-2006 ETDs are digitized and made available if proper permission and approvals are secured. ETDs are additionally discoverable via our catalog through an ETL process which transforms CDR metadata to the format used by our discovery layer. Recently, we have been in discussion with the Graduate School at UNC about the approaching ADA Title II accessibility deadline in order to implement accessibility policies and workflows into the ETD submission process.
Additional resources:
Flynn, E. A., & Ahrberg, J. H. (2020). Electronic Theses and Dissertations (ETDs) Metadata Policies, Workflows, and Practices: A Survey of the ETD Metadata Lifecycle at United States Academic Institutions. Journal of Library Metadata, 20(2–3), 91–110. https://doi.org/10.1080/19386389.2020.1780689
Hollingsworth, C. (2025). Creating New Access Points for Print Theses and Dissertations. Repository@TWU. https://hdl.handle.net/11274/17287
Lake, S., & Nicholson, J. (2023). Remediation by Degrees: Enhancing ETD Metadata to Improve Discoverability. North Carolina Libraries, 81(1), 13. https://doi.org/10.3776/ncl.v81i1.5426
Thompson, S., Liu, X., Duran, A., & Washington, A. M. (2019). A Case Study of ETD Metadata Remediation at the University of Houston Libraries. Library Resources & Technical Services, 63(1). https://journals.ala.org/index.php/lrts/article/view/6764/9320
tl;dr btrix is a
bring-your-own-model chat interface for web archiving, which I demo
below with my shiny new inference.coop account
I’ve been experimenting
recently with a chat interface to the browsertrix-crawler
software. If you haven’t seen it before, browsertrix-crawler is an
excellent tool for creating high-fidelity archival snapshots of web
pages and websites. It has many, many options, and the browsertrix service provides a full
user interface to managing one or many browsertrix-crawler jobs.
So why would you ever want to operate browsertrix-crawler with a chat
interface?
I guess the main reason is that you may feel like its overkill to bring
up and run a browsertrix instance for the type of job you are
doing. Another is that perhaps you would like to explore the options for
archiving a website, including developing a custom
behavior for clicking on parts of the pages. Its possible that the
chat interface provides a more iterative way to discover the many
affordances that browsertrix-crawler makes available for web archiving,
in the context of actual web archiving work.
Ok, and truth be told, it was an opportunity to experiment with how well
this approach could work. I think the answer is mostly, yes.
The seed for this idea was a set of simple shell scripts, dubbed browsertricky, that I’d
created a few years ago for running browsertrix-crawler jobs. It helped
me keep around the YAML configuration files in a single location for
reuse, and took some of the guesswork out of remembering how to
configure browsertrix-crawler, and monitor its progress. This sort of
interaction on the command line was something I was interested in seeing
in the chat style interface.
I initially developed what became btrix as a Claude plugin to help
explore how to control browsertrix-crawler with Claude Code. But after
getting it mostly “working” I wasn’t really happy with the broad
permissions it needed to be granted to operate and introspect on
browsertrix-crawler. I found that Claude improvised too much, sometimes
reaching to write and run its own programs to work with
browsertrix-crawler, instead of using the skills and scripts I had
explicitly built for it to use. I also didn’t really like the lock-in to
Anthropic, or any single LLM model.
So, long story short, I ended up developing btrix using the pi framework, which lets you connect to a
large number of backend models. In the demo below I use
deepseek-v4-1-flash via Inference Cooperative, which is Nathan Schneider’s latest joint
(more about this in the future). pi let me write tools for the
framework to call out to. Tools have known programmatic interfaces that
can be tested, and pi can lock down other types of interaction.
So how does it work?
Below is a short-ish demonstration of using btrix to archive a small
part of the news.stanford.edu website. It shows how btrix is able to
write a custom behavior based on my description. btrix writes the
JavaScript behavior, the YAML configuration, and starts the crawl. It
provides diagnostic information about the crawl as it progresses. It
also shows a quick look at the artifacts it leaves on the file-system,
including the WACZ archive, which it dutifully plays back.
As you will see, the session wasn’t without its hiccups because rather
than using a big corporate LLM I chose to use my new Inference
Cooperative account, which is really brand new, and seemed to encounter
a brief outtage after the crawl completed. But it recovered!
In case you missed it, here’s what btrix wrote at the end:
The archive is being served now — open this in your browser:
http://127.0.0.1:8087/?source=stanford-news.wacz
No notes came back with it, so there's nothing to warn about; it's served from
the same origin as the archive, so replay needs no network access.
What you should see: the listing page, and the paging buttons should still work
— clicking 2 and 3 replays the archived API responses rather than hitting the
live site. The 30 stories are all in the WACZ as their own pages, so full-text
search in the replay UI should find them.
For reference, the behavior did what it was supposed to, per the crawl log:
- page 1 — 10 queued
- page 2 — 20 queued
- page 3 — 30 queued
All 31 pages returned HTTP 200, median page text ~11.9k characters. Worth
a look while you're in there: the numbered buttons go up to 5, and I only
clicked through to page 3 as you asked. If you'd like pages 4 and 5 too, it's
a one-line change to the behavior and a re-run.
One of the handy things that btrix does is embed its own replay server
which uses replayweb.page to let you inspect the WACZ file you created
with ReplayWeb.page. If you get a
chance to give it a try please let me
know how it works for you.
Lately I've been playing classic World of Warcraft to get a feel for it and see if I like it before I try and go head-first into WoW Forever fairly blind with some friends. People have told me that World of Warcraft is good, but I haven't personally played it (or really any MMO outside of Final Fantasy 14). I've wanted to see for myself, so I dug in.
I've been leveling an undead warlock and recently I've gotten my way to The Barrens. My goal was to kill some planestriders and collect their beaks as proof of my kills so that the population can get back under control or whatever. I wandered around and found another player also on that quest and waved to them. He was a gnome warlock named Borton and waved at me when I walked over. I waved back and invited him to a party. We walked around the area and started spreading our aggro to farm the planestriders and grab their drops. After a few kills he stopped for a second and said "this is why we play".
I was taken a bit aback by this. It felt out of nowhere but also felt like one of the more human encounters I've had with people in a video game in a long time. One of the great parts about MMOs is that you end up having a lot of these itinerant bits of community like this. You both get a quest from an NPC asking you to go and farm some drops and then this sends you all out into the world to naturally do it together. These conversations and liminal bits of community end up adding up into the greater experiences and really giving you the flavour of the game's vibe. You do it together for the fun of the game.
After a while we finished our farming, got all the planestrider beaks, commiserated about the drop tables, and decided to just sit down on a nearby hill and take a screenshot because it just felt right.
Two players sitting in golden fields with a backdrop of mountains and a blue sky enjoying the moment together wordlessly.
This has really gotten me thinking about what I really want to get out of games and why I play them. I think I've been doing it wrong.
The endless treadmill of “dead content”
At some level the “high end” content of online games like Final Fantasy 14 or Retail World of Warcraft looks endless. There’s intricate dances of choreography as you and your friends puzzle your way through each boss mechanic towards victory. Hell, games like Final Fantasy 14 have their fights defined by the little call->response pairs as the boss telegraphs things and you react to it in concert. Not to mention the feeling of utter serotonin that comes from figuring it out and winning.
One of the main problems with this is that the “high end” raids in this game tend to be a mile wide in theory but practically an inch deep. Don’t get me wrong, even as far back as a decade ago there’s still encounters that are actually hard (eg: XIV’s ultimate raids, I'm personally at Quickmarch Trio in UCOB if you have the right brainworms to know what that means), but as the fights get older there’s a sharp dropoff in the number of people that actually want to run them. This ends up with entire subsystems of the game (Deep Dungeons, Bozja before Shadowbringers went into the free trial, etc) becoming “dead content” because if you queue for them, nobody shows up. If you put up an entry for them in Party Finder, nobody will join you unless you set aside time to wait for hours. You end up with modes where you're encouraged to do it in a group being bereft of people so you're "forced" to figure it out on your own.
As a result everyone is funneled into the high end and trying to min-max it as much as possible to squeeze all the game they can out as fast as they can before “everyone else” gets what they want out of it and moves on. The world becomes a theme park where only the newest rides have the most people there and getting a group for anything else becomes a slog in party finder or needing you to seek out people that focus on that thing and nothing else.
This also means that if you’re “slow” or have to take a break during the small sliver when content is “current”, you’re going to be stuck "behind" on it for a long time without significant effort to find a way around it. It really makes you wonder what the point of it all is if you can’t find anyone to do it. Worse, when you're trying to find a group to do it regularly (a "static" group), they look at how you did on previous raid tiers as a way to vet if you're a "shitter" or not.
The records of your ultimate achievements of an entire raid tier, or the relic weapons for an entire expansion just end up becoming things you can do if you want that give you a shiny skin for your weapons to coordinate with an outfit that looks better. Maybe in edge cases those stats will be better for complicated maths reasons. All that effort you put into them is just ultimately "wasted" when the next ride comes out, all the stats get reset, and everyone moves on to wait in line for the next thing. All of your gear you spent weeks grinding for is reduced to sticks that increase the statistics correlated to aspects of your character that make the funny numbers go up so you can kill the boss faster to move on to the next thing.
Maybe this is just something endemic to the theme park design philosophy where you’re given a goal by the developers and a series of actions that really boil down to TODO items that are the barrier from you and that thing you actually want (do these quests, fight these normal raid bosses, beat the savage versions of those bosses, beat the ultimate boss, get the shiny weapons, and finally look cool afk-ing in Limsa Lominsa). Maybe this is why people end up seeing things like the overworld as the barrier to the rides instead of an integral part of the experience. Maybe this is the pressure that makes people want to have flying mounts or other ways to "skip the grind".
A meditation on Death Stranding
I haven't really touched Death Stranding since I did my review on it way back in 2019. It's always had a special place in my heart though, it's kind of a touching experience even though it's incredibly isolating at the same time. This game has two main characters: Sam "Porter" Bridges and the game world. The overworld's system of traversal is the protagonist as much as the actual protagonist is.
Sam "Porter" Bridges looking out past the hill he just climbed and seeing an upside-down rainbow.
God sent the rainbow as a promise to mankind that he wouldn't try to destroy them again, and this rainbow is upside down after mankind's near destruction.
Honestly I think one of the best features of Death Stranding is the camera mode. At any point you can hit a button (on the Steam Controller it's one of the back buttons by default) and the action stops so you can position the camera and just soak in the scenery. If you've never played Death Stranding before, I highly suggest you stop reading this post right now and buy it on whatever platform you buy games on. Worst case you can experience it on the iPhone as Hideo Kojima surely intended.
If you actually do this on an iPhone, please do it with a controller. The
touch controls are kinda awful.
As you progress your way delivering packages to save America, every so often the camera will pull back, the game actions will become muted, and the normally contemplative and silent experience will suddenly be augmented by soft music.
The protagonist of Death Stranding riding on a tricycle down to a port city with the Death Stranding logo overlaid.
It turns this process of progression into a serene bit of negative space in a way that makes you really stop and think about the experience, the dramatis personae, and the kind of story that's trying to be conveyed. The world itself turns into this sense of wanderlust that makes you mentally pause and just process things.
I don't really know if I'm doing this justice, in the moment it's an almost spiritual kind of connection between you and the overworld mediated by the soft sounds of Silent Poets' Asylum For The Feeling or Low Roar's Don't Be So Serious. The game makes you feel so small yet so connected to the world around you and there's just this overwhelming sense of equanimity.
Just take a moment to watch the Don't Be So Serious video with Death Stranding gameplay underlaid. That's a snapshot of how the game literally plays out. It's these liminal bits of wanderlust combined with a story that only Kojima or Yoko Taro could come up with.
I miss the feeling of exploring
This kind of overwhelming emotion is the kind of thing that made me fall in love with gaming back in the days when the most complicated games I'd play would be Banjo Kazooie, Super Mario 64, Super Mario Sunshine, Pikmin 2, Sonic Adventure 2, and other games of that vintage. At some level I kind of long for not knowing what's beyond the next hill and being forced to figure it out by just going there to find out.
Hell, I also get this sense of romance in Breath of the Wild or Tears of the Kingdom when I just turn off all the UI elements and wander around wherever the wind takes me.
Link in Breath of the Wild looking towards a sunset bathing a lake in twilight as the light refracts on the smooth stone surface he's standing on.
Breath of the Wild is one of those games where you can just wander around the overworld for ages and still find new things. Combine it with flying and you just feel things as you go from hither to thither. Even with fast travel you still have to go there yourself. You have to climb every mountain. You have to cross the vast expanse of the desert. You have to figure out how to get a boat to the island with some awesome gear. And even months or years after you capture screenshots/video of the adventure, looking at them takes you back there and you just feel it again.
You have the power and you have to pick what you do with it. There is no wrong way to do it. Want to fly over Gerudo Town after sunset knowing well and good that the moment you land you're gonna get kicked out of it for being a guy? Hell yeah, you can just DO IT.
Link in Breath of the Wild gliding above Gerudo Town at sunset, just soaking in the vibes.
This is making me think, fuck, am I actually getting what I want out of my time? In my race to min-max my way through "current" things so I don't get "behind" have I just lost that sense of wanderlust from the experience? I miss that feeling of just exploring the world for what it is and letting myself get there the long way.
Have I been skipping over the real "good part" in my effort to be efficient in getting to what I thought the "good part" really was?
Have I just been playing games the wrong way?
This is why we play
I think I've been playing the games I'm immersing myself in incorrectly and I want to change that. I want to lose myself in the fantasy worlds. I want to let myself see the sights, hear the sounds, feel the breeze on my back, go fishing at what feels like the little corner of civilization just at the edge of the universe, and just let that sense of overwhelming wanderlust come back into my life.
This is why we play, we play to let the fantasies of these fantasy worlds take charge and show you something that you didn't think you needed in your life yet you're all the richer for it. As World of Warcraft Forever comes out in November and Final Fantasy 14's Evercold expansion comes out next year, I'm going to just let those worlds enrapture me in their designs and lose myself to the wanderlust. I'll let The Fourth take me where it does and that's okay.
It's absurd that the root of this revelation was a bot in a single-player MMO on my homelab picking a random chat message out of a list of dozens randomly triggered by a random overworld mob dying.
Yeah, I forgot to mention, this all was caused by an experiment in playing with bots to teach me what it means to have a community. I was gonna lead this post with that fact but I think it all lands better if I just spill it all out at the end while I'm having the revelations like this. Given the utter absurdities of this modern era where we live in Clown World, I think it only makes sense that it would boil down to something like that.
Honestly I think it's even more absurd that it took a proc of a proc of a bot saying things for me to figure out what was wrong with myself.
I feel like I've just torn out a chunk of my soul and rended it unto the canvas for this. I don't really know why this helped me, but I hope this helps you too. Stay fresh out there, and if you're going to also venture forth to that millennial retirement home of a launch, maybe I'll see you in Azeroth in early November for World of Warcraft Forever. I'm gonna go in blind, you should too. We'll figure it all out together, one terrible loot table at a time.
The guest editors for this special issue showcasing examples and ideas for integrating GenAI tools into galleries, libraries, archives, and museums generative AI summarize the articles included in it.
“Prompting Critical Thinkers” follows one librarian’s evolution of a library lesson on generative AI tools into an information literacy lesson grounded in the ACRL Framework for Information Literacy for Higher Education. The revised lesson focused on generative AI as an information source and blended definitions of AI literacy from academia and industry. This case study demonstrates how growing interest in AI skills can give librarians new opportunities to get into the classroom and teach students enduring information literacy skills.
Large-scale historical corpora present persistent challenges for research, particularly in relation to the navigation, interpretation, and extraction of meaningful information across extensive and heterogeneous collections. The Legislación Mexicana corpus, a 42-volume compilation of legal dispositions spanning more than two centuries, exemplifies these challenges due to its size, structural complexity, and evolving terminology. Traditional approaches to working with such corpora have relied on close reading and the use of indexes. While effective within defined limits, these methods constrain the scope of inquiry and require significant time and expertise to produce meaningful results. This article presents the development of LegMexIA, an AI-enabled research assistant designed to extend interaction with the corpus beyond conventional retrieval methods. The system integrates traditional text retrieval and retrieval-augmented generation via a coordinated architecture that distinguishes between different types of user queries. Structured queries are processed using Elasticsearch and BM25 ranking, while exploratory and interpretive queries are routed through a retrieval-augmented generation workflow, where relevant text fragments are retrieved using vector similarity and assembled into contextual inputs for a language model. A key aspect of the system is the agentic decision layer that analyzes user intent and dynamically selects the appropriate processing strategy, in alignment with established reference practices in academic libraries and in a way that maintains the user’s role in verification. The article also addresses key considerations related to prompt design, technological sustainability, institutional constraints, and the implications of public deployment.
This paper reviews Project Gutenberg’s (PG) use of artificial intelligence (AI) with the free books in its digital library. The paper covers three main projects where PG has integrated AI into its workflow. The three areas discussed are disparate in nature and include audiobook, book summary, and browsing category generation. The paper reviews the AI models and processes used, as well as how volunteers and users have responded to the organization’s use of AI. PG’s dataset is extremely large and includes more than 75,000 books. Large datasets can be challenging and costly for AI models to process. The main goal of this paper is to share how PG has incorporated AI processes into the workflow of a very large digital library while maintaining human oversight.
Many heritage institutions hold extensive collections of show programs. Too numerous to be cataloged individually, they remain difficult to access and are still largely undigitized. This paper presents a workflow for transforming such documents into structured, interoperable data by combining vision-language models (VLMs), a custom extension of the Linked Art ontology for the performing arts, and different approaches for automatic semantic data annotation. Using the Festival d’Avignon programs (1947–present), preserved at the Bibliothèque nationale de France, as a case study, we demonstrate how VLMs achieve more than 98% text transcription accuracy from heterogeneous heritage document images and how constraining large language model generation through a knowledge graph significantly reduces hallucinations in semantic annotation. We argue that this workflow opens new possibilities for large-scale, interoperable performing arts historiography, including the creation of catalogues raisonnés for performing artists, a form of scholarship common in visual arts but absent in performing arts studies.
This paper explores the emergence of artificial research intelligence (ARI), autonomous AI systems designed to conduct end-to-end scientific research, from ideation to manuscript production. Unlike simple AI enhancements, ARI utilizes ensembles of agentic models to automate the scientific method, potentially creating a “fifth scientific research paradigm.” While ARI offers rapid, low-cost discovery, it faces significant challenges including “epistemic capture,” where knowledge production shifts from public universities to private AI corporations. Technically, these systems currently struggle with hallucinations, a lack of deep domain expertise, and a tendency to overclaim results. However, the growing number of ARI frameworks points to an emerging research methodology of transformational possibilities. For research libraries, ARI necessitates a profound shift from managing static collections to stewarding dynamic, AI-driven knowledge systems. Libraries can lead by reimagining processes and services supporting scholarly communications, enhancing AI competencies, and conducting applied research through library-based AI labs. The combination of advanced compute and the guiding framework of the scientific method makes ARI a human–AI co-creation model where AI performs the "reckoning" (computation) and humans provide the "judgment" (ethical and intellectual oversight).
The rapid integration of generative artificial intelligence (GenAI) into academic library discovery systems offers new opportunities but also introduces challenges, as vendor tools streamline natural language searching while obscuring search logic. This article presents a case study of a librarian-designed AI discovery tool that is focused on transparency and pedagogy. The assistant converts natural language queries into structured, editable Boolean search strings; surfaces search logic; and retrieves results across library holdings. Developed through iterative prompt engineering, this tool demonstrates how GenAI can support research practices without sacrificing rigor. Findings from qualitative feedback and quantitative assessment indicate that the tool improves the consistency and quality of Boolean query construction while fostering user engagement with search as a reflective process. This case study argues that libraries can play a critical role in shaping AI discovery tools that align with information literacy goals and institutional scholarly standards. By foregrounding transparency and user agency, librarian-led AI design positions GenAI as a form of instructional scaffolding rather than a replacement for expertise.
Academic libraries face ongoing difficulty in aligning their collections with the needs of the communities they serve. Although the literature identified the value of using research, publications, and teaching data to inform collection development, maintaining crosswalks between these activities and library classification schemas made such work at scale impractical. This study developed a scalable workflow that maps institutional outputs to Library of Congress (LC) classification categories using a large language model to support an ongoing, evidence-based approval plan review. The workflow integrated six datasets across two domains of academic activity: research outputs and teaching activities. Where bibliographic metadata existed, they were used directly; otherwise, Anthropic’s Claude Sonnet inferred LC classifications through data-specific prompts. A Python pipeline combining the Claude and bibliographic metadata APIs produced enriched datasets at scale, followed by automated LC range validation and selective human review. The outputs feed a three-page interactive Tableau dashboard: a summary landing page, a Research Outputs Explorer, and a Teaching Activities Explorer. The collection development team has used the dashboard alongside expenditure and usage data to review approval plans, highlighting the value of unifying previously siloed data sources. The workflow is designed to refresh with new outputs and improved models, offering a sustainable foundation for ongoing collection assessment, including gap analysis and alignment with institutional priorities.
Amazon Web Services cannot restore access to its cloud-computing
facility in Bahrain and one of three data-hosting zones in the United
Arab Emirates following damage during the Iran war, according to a
status update seen by Reuters on Tuesday.
The idea of Kantian ethics is both simple and revolutionary: it proposes
a moral law independent of any notion of a pre-established Good or any
‘human inclination’ such as love, sympathy or fear. In attempting to
interpret such a revolutionary proposition in a more ‘humane’ light, and
to turn Kant into our contemporary—someone who can help us with our own
ethical dilemmas—many Kantian scholars have glossed over its apparent
paradoxes and impossible claims. This book is concerned with doing
exactly the opposite. Kant, thank God, is not our contemporary; he
stands against the grain of our times. Lacan on the face of it appears
the very antithesis of Kant—the wild theorist of psychoanalysis compared
to the sober Enlightenment thinker. His concept of the Real, however,
provides perhaps the most useful backdrop to this new interpretation of
Kantian ethics. Constantly juxtaposing her readings of the two
philosophers. Alenka Zupancic summons up an ‘ethics of the Real’, and
clears the ground for a radical restoration of the disruptive element in
ethics.
And just because AI won’t eventuate in a god-device, even one that goes
rogue and unleashes biological warfare on humanity, that doesn’t mean
that LLMs, particularly those with safety standards relaxed in order to
compete, won’t behave in unpredictable and potentially ruinous ways. Nor
does it mean they won’t be used by hackers, criminal gangs,
dark-money-funded political campaigns, blackmailers, and various others
to sabotage public bodies and infrastructure. Just as there are many
problems with the advertising industry that don’t involve subliminal
mind control, so the problems with AI that don’t involve a god-device
turning against humanity are legion.
For weeks, Washington policymakers have been intensely debating AI after
a series of dire warnings from Silicon Valley engineers and tech CEOs of
the possibility that AI could break free of human constraints, with
potentially civilization-ending consequences.
But the episode with the Chinese ship underscores a different, and more
immediate risk: human beings making disastrous decisions based on
inaccurate or misleading information generated by AI or other automated
systems. Sources said that the military is rapidly turning to AI to help
with targeting, an area which holds the obvious risk of fatal mistakes.
I’ve always had a love/hate relationship with string as a software
engineer, but this is slightly different. Take a support system that
needs to know whether a customer is asking for a refund. With a
generative model the application asks a question, gets back text or
JSON, validates it, and turns it into a branch in the program.
Structured outputs make that arrangement less painful, but the interface
underneath is still generative. The application needs a decision. The
model produces tokens that represent one, and the application has to
trust the representation. .
Jev starts from the decision space instead. The API exposes three
primitives: Choice picks from declared alternatives and returns their
probabilities plus a confidence value. Score evaluates ordered
descriptive levels and returns a continuous score, a distribution, and
confidence. Noul evaluates a binary proposition and returns the
probability that it is true. You can combine all three in one request
against the same state.
Even the most complex software is built out of simple logic and layered
abstractions, with every branch auditable. We want AI to work alongside
existing software as a primitive that any programmer can invoke for
semantic judgement and decisions, while still using code for what it’s
best at: exact computation.
Trust in U.S. science policy has fallen to a low point in Europe, and
fears of losing scientific data are mounting. The U.S. government quit
both the U.N. Intergovernmental Panel on Climate Change (IPCC) and the
U.N. Framework Convention on Climate Change in January, though U.S.
researchers are still contributing to global climate science. To support
them, the German Research Foundation in late 2025 set up a program, with
a budget of $35 million, to provide a safe haven for permanently storing
copies of U.S. data from public and open sources.
Why did entrepreneurs in early aviation converge on the same purpose in
each period (entertainment, then military production, then airmail, then
passenger transport), even as they brought diverse backgrounds and
experimented vigorously over how to pursue it? We develop the
Systems-Based View (SBV), a grammar for analyzing how socio-technical
systems structure entrepreneurial action. SBV is built on four
primitives: purpose, artifacts as systems, systems as technology, and
bottlenecks. Drawing on a deep historical analysis of U.S. aviation from
1903 to 1937, we show that the evolving socio-technical system
constrained feasible purposes to a narrow set in each period, while
permitting abundant experimentation within them, and that the resolution
of successive bottlenecks expanded the set of feasible purposes over
time. SBV reduces the dimensionality of strategic analysis by directing
attention to three questions: what is feasible given current
constraints, what bottleneck is binding, and where do actors’ maps of
the system diverge from its actual structure? We specify boundary
conditions for when SBV’s constraints are most likely to bind, rank
empirical tests by their vulnerability to alternative explanations, and
show how SBV complements existing strategic frameworks by identifying
the constraint landscape within which they operate. SBV is calibrated on
early aviation; its value will depend on how well it travels to new
settings
Thanks to the September release team: Martha Driscoll (NOBLE), Blake Graham-Henderson (MOBIUS) Gina Monti (Bibliomation), and Andrea Buntz Neiman (Equinox); as well as everyone who contributed fixes and testing to this release.
The Worry Box in the Wellbeing Room in the Radcliffe Science Library, Oxford. Students drop in notes about what’s worrying them; the library shreds these notes and they become fertilizer for new plants at the Oxford Botanic Garden.
* * *
One of the most damning descriptions of Lyndon Johnson in Robert Caro’s The Path to Power is from an ostensible ally of LBJ’s, who said that Johnson “listened at” other people, rather than to them. It feels like our modern media environment encourages all of us to do the same. As a modest corrective to this inclination, I’m adding occasional posts with reader feedback and writing, highlighting voices other than my own and related ideas I believe are worth listening to.
In “Phantom Limbs, Old and New,” an insightful response to the last piece in my series on AI and scholarship, which addressed the preservation of the scholarly record, my friend and frequent collaborator Tom Scheinfeldt, a professor at the University of Connecticut, makes the excellent point that important parts of that record have always been difficult to store and retrieve:
The emergence of large data sets and LLMs reveals a problem we’ve always had but never needed to face. Scholarly communication has never been a true mirror of scholarly work. There have always been phantom limbs in the research process, the most significant of which are uncredited collaborators: wives, students and other “invisible technicians” in Steven Shapin’s phrase. There are also tacit processes, laboratory habits, and assumed knowledge that lurk behind and beneath published work that go undescribed and unquestioned. In between what goes into research and what comes out of it is a hazy space filled with research assistants and tacit knowledge. The body of knowledge has always had phantom limbs, and it always will.
On the nondeterministic nature of LLMs, which I noted makes them elusive as trustworthy sources of scholarly analysis, Tom counters:
It is true that, even if we could somehow preserve today’s LLMs, their nondeterministic nature means re-running them would produce slightly different results each time. But it’s also true, as Otto Sibum showed, that Joule’s contemporaries could not reproduce his mechanical equivalent of heat experiments because they lacked the manual skills he had learned as a brewer. And a retired model is no more inaccessible than a dead scientist. Ultimately, we rely on whatever the output communicates.
Tom’s reminder of the role of tacit knowledge in so much of experimental science is also an implicit criticism of the AI boosterism maintaining that AI + lab automation = a cure for cancer. There are often distinctive techniques, elements of craft, that aren’t captured in published articles. Up-and-coming labs are excited to hire scientists from leading labs not just because of their knowledge, but because of all of the little, sometimes unarticulated or even inexpressible, ways that those leading labs do things, which can make all the difference between success and failure.
After my post on “Vibe Analysis,” or using AI to create small exploratory environments or visualizations that aid in the development of a scholarly thesis, my Northeastern University history department colleague Chris Parsons shared a remarkable website, Relations des Jésuites de la Nouvelle-France · 1632–1672, that he created with Claude to help with his new book project on the French colonization of what is now Canada. Chris downloaded digitized editions of all of the still-extant assembled reports that the Jesuit mission in “New France” sent back to Paris. Running locally on a Raspberry Pi (!), the site is enormously useful as a tool to advance Chris’s scholarship, and as an aid to other scholars in the history of colonial North America.
With terrific affordances that allow visitors to go back and forth between the original books, the OCRed text, and modernized versions of the French using resources from the open-source FreEM project (the original French, as with English from the early modern period, has variations in spelling and grammar that sometimes make it hard to read), this is a great example of what a single scholar can produce with a little AI, a lot of knowledge about the resources in a field, and open-access library collections (in this case, including volumes from the John Carter Brown Library, BnF Gallica, and the Thomas Fisher Rare Book Library at the University of Toronto).
Page image with transcriptions in the original early modern French and modernized French, Relations des Jésuites de la Nouvelle-France · 1632–1672
Chris has also been able to use Claude to create an index of the people mentioned in the texts, including, for the first time, a thorough concordance of the Indigenous Wendat people the French encountered. Additional data visualizations provide other pathways into the rich texts, but the texts themselves are never abstracted away — they remain right at your fingertips.
This is precisely what I’ve been trying to imagine in my AI and scholarship series: a combination of close reading and AI tools that help the scholar with deep rather than superficial research. The idea of a small virtual bookshelf as a bespoke scholarly resource, supplemented by LLMs and other digital methods, also presents itself as a good use case for our new Mellon grant on AI and books.
Grant wrote to me in response to that piece to highlight the ML.ENERGY project, led by the University of Michigan. The collaborators on this project test machine learning routines and other forms of AI on different GPUs, and with a range of models, to assess how much energy is used by each unit of a task. They regularly publish leaderboards and associated charts that can inform ecologically sensitive usage of this technology.
ML.ENERGY dashboard for joules expended per token for text production, for a range of open-weight models
What one immediately notices looking at these leaderboards is that AI energy usage varies extremely widely, with some model-task pairs using a tiny fraction of the energy that other model-task pairs use. In the example above, in which an LLM was asked to produce text as part of a conversation, the worst performers used over 10 joules per token (roughly, a word), while the best used well under a joule per token, and produced the text just as quickly as the energy hogs. So with some attention to the task, model, and computing environment, it seems very possible to use AI with as little as 1 or 2% of the energy you would expend if you threw the same task mindlessly at whatever model you normally use.
If a library like mine used AI on a large corpus — millions of documents or images — this gap would be multiplied many times over, and the energy savings would be tremendous. Currently the tasks tested by ML.ENERGY are common ones, like producing a stream of text; we will need more tailored task assessments for different academic disciplines, such as translation, handwriting transcription, and photographic analysis. But there is a pathway forward here to sharply reduce the environmental impact of AI — if we choose to take it.
Finally, my Northeastern University Library colleague Lawrence Evalyn, our Text Mining Specialist, has also been exploring how to use AI to supplement and assist, rather than replace, human scholarship. He wrote to me with a wonderful personal case study:
I am interested in the eighteenth-century practice of publishing “by subscription,” which was essentially a form of crowdfunding akin to a modern Kickstarter, complete with the practice of thanking one’s supporters in a published list. A few thousand books were published by subscription in England prior to the nineteenth century, each with names of a hundred to a thousand individuals: a tantalizing information source to study at scale, but also a challenging one to make computationally tractable. From the 1970s to 1990s, the Book Subscription Lists Project at the University of Newcastle upon Tyne undertook to collect and “computerise” 8,330 subscription lists from 1680 to 1794. This was an amazing feat of bibliography! The project published bibliographies as printed books, but as far as I can tell, no data has been preserved digitally.
Five years ago, this is where I would have given up: it would be possible to re-transcribe all of the bibliographies, but I would never choose to go down that path. I would have cried to a friend about the fragility of magnetic tapes as a storage medium, and tried to think of a smaller, easier project.
Last week, though, I scanned a few pages just to see, and within about three hours of working with Claude Code I had a pipeline which pristinely takes those page images and makes me a spreadsheet of all the bibliographic data and metadata. Just look!
The original; like Proust’s madeleine, this font evokes memoriesTranscribed and put into rigorous tabular form; nice job, Claude
I especially appreciate Lawrence’s marriage of cutting-edge techniques with the continued importance of print and the library (and, critically, interlibrary loan). His conclusion, in a note to me:
I’m thrilled by the possibilities this opens up for my research, and pleased by this example of LLMs working in concert with “traditional” scholarship...and a bit pensive about how I can make something that would be usable for the next fifty-years-later scholar.
Recently, interest rates for government debt have been rising worldwide. As I write the yield of the 10-year US treasury is over 5%. It had been below 4.5% for over a year when, in March, it started a steady rise. This is despite (and probably because of) Scott Bessent's efforts to force it lower, which JP Morgan compared to 'paying your mortgage with your credit card'.
A big part of the reason is governments persistently borrowing to cover deficits, with no plan to reduce them. For example, the Trump administration has borrowed more than the growth in US GDP. Projections are that by the end of their term they will have borrowed $1.40 for every $1 growth in GDP. Half of that GDP growth is from the AI bubble. Remove it and the administration has already borrowed almost 2.5 times the GDP growth.
But another part of the reason is that governments are facing competition for the available funds from other highly-rated borrowers. Hyperscalers such as Microsoft (AAA) and Alphabet (AA+) can no longer pay cash for the immense sums they believe they need in order to keep up with their competitors in the race for AGI. So the bond market is facing a significant increase in demand, which naturally increases interest rates.
Below the fold I look into how this competition for funds is going.
It is a fact of life that the more debt you carry, the worse your credit rating and thus the more expensive it is to add to your debt. This is why the hyperscalers are so anxious to have their massive borrowings kept off their balance sheets and attributed instead to "Special Purpose Vehicles". They hope in this way to distract the rating agencies enough to prevent being downgraded. A downgrade would increase their interest expenditures, thus reduce the coverage of their interest payments, and potentially lead to a further downgrade and a death spiral.
Achieving an investment-grade rating from Fitch, Moody’s and S&P soon after going public would be a remarkable feat for the two lossmaking AI labs, unlocking big benefits for the companies and their infrastructure partners including Oracle and Nvidia. …
Analysts at rating agencies are waiting to see the results of their IPOs before reaching a decision. The two companies remain unprofitable and have shown little sign of generating positive free cash flow. They also face growing risks, including the popularity of Chinese open-weight models.
“We still treat OpenAI and Anthropic as deep in speculative grade . . . they are in the red,” said another senior credit analyst.
An investment-grade rating of BBB- or above would allow them to borrow from institutional lenders such as pension funds.
S&P reckons a full half of the economic growth coming from the US private sector was linked to AI-centric activity over the past year. And they assume the top six US hyperscalers will collectively spend more than $7tn in the five years through 2030. So the fate of the US economy, not to mention the stock market, credit, infrastructure and real estate, is increasingly tied to the AI show staying on the road.
Every time we take a deep dive into this sector, we find that capex is rising faster than we anticipated, financings are becoming more complicated and less transparent, and that returns on investment will take years to realize.
The credit rating process involves analysts diving deep into the workings of companies, typically with access to a ton of non-public information. But the authors write that, despite AI return on investment being critical to the rating judgments, “the big six hyperscalers don’t quantify their returns on investment on AI”. Maybe, the report’s authors speculate, this is because it’s difficult to work out. Or maybe, “simply, entities just choose not to share that data.”
The reason these companies "choose not to share" their "return on investment on AI" with the rating agencies is left as an exercise for the reader.
Source
Nangle has a set of interesting charts based on S&P data. This one shows the ratio of debt and "debtlike commitments" to EBITDA (earnings before interest, taxes, depreciation, and amortization) for the six largest US hyperscalers. Each company is labeled with their current rating. Horizontal dotted lines indicate the trigger level for this ratio that could cause a downgrade. Note that:
Oracle's projected 2027 and 2028 commitments are so close to the trigger for downgrading from their BBB- rating that even a small drop in earnings would render their bonds junk. This would force institutional investors to sell them, making future borrowing and rolling over existing debt as it comes due much more expensive.
Amazon's projected 2027 and 2028 commitments look as though they would cause a downgrade from AA, but it is more likely that Amazon can increase earnings fast enough to prevent this.
Alphabet has a lot of headroom before their AA+ rating is at risk.
Source
An alternative way of looking at the situation is to ask how much more could these companies borrow, or how could their EBITDA decrease, each year before they would be downgraded. This chart answers the question. Note that:
For 2027 through 2029, only Alphabet can borrow more than $100B each year.
On the current projections Amazon will, and Oracle is extremely likely to, be downgraded.
Nangle makes an English understatement:
For double-A-rated Amazon to move to high single-A wouldn’t be much of a disaster. But in Oracle’s case, getting downgraded means getting junked. And with around $117bn of index-eligible US dollar bonds, that would be quite the event.
The limited headroom for increased borrowing starts next year, which is when the "take or pay" commitments with data centers start costing serious money.
Source
The third chart provides yet another way to look at the situation. It plots the companies' projections for EBITDA against the threshhold for downgrades (dotted line). Note that:
Amazon has to exceed and Oracle has to meet their EBITDA projections if they are not to be downgraded.
All these companies' projections for their EBITDA growth from 2025 to 2029 are astonishing. Eyeballing the chart we have:
Alphabet: 2.5x
Amazon: 2x
Meta: 2.6x
Microsoft: 1.8x
Oracle: 3.5x
SpaceX: 15x
Nangle's comment that these represent "pretty punchy EBITDA growth" is another understatement. For example, Alphabet is saying it will add around $225B in EBITDA by 2029. I'm an Englishman, so I can say that it isn't clear where an additional $225B would come from.
Alphabet and Meta together are in effect projecting that by 2029 companies will increase their advertising budgets by $385B/year. Of course, this isn't going to happen. All these companies are expecting the EBITDA growth to come from the "returns on investment on AI" that they "choose not to share" with the rating agencies.
Ignoring the ludicrous SpaceX projection, the other 5 companies' "punchy" projections assume 2029 increased EBITDA of around $900B, which is less than half the estimate of over $2T/year in additional revenue needed to cover their borrowings.
Is it plausible that in 2029 OpenAI and Anthropic will together have more than $1.1T in revenue to make up the difference?
The report also features a circular financing taxonomy, which includes residual-value guarantees, take-or-pay agreements, chip financing, lease liabilities, power purchase agreements, backstop guarantees, lease guarantees, and direct equity investments.
interconnected financing structures could amplify volatility if demand weakens unexpectedly. The failure of one entity would have implications for entities that have lent it money, have guaranteed the value of its assets, or are just expecting payment for goods delivered.
The broader question is whether circular financing is creating leverage collectively that is individually manageable but could become highly correlated should demand fall sharply. . . .
The scale of overlap and interconnectedness is vast. In a downturn, even the best capitalized and most profitable firms may incur substantial pain.
Ed Zitron's Concentration Risk points out another big risk to the hyperscalers' EBITDA projections that S&P seems to have missed:
Anthropic and OpenAI’s compute commitments, in my mind, should be seen more as debt obligations than “contracts,” because they (as take-or-pay agreements) function in much the same way, requiring the company to pay whether or not they need the capacity.
For now, everything looks awesome. Microsoft, Google and Amazon have all had big bumps in revenue from AI lab compute spend along with massive, ever-swelling revenue backlogs — over $1.5 trillion worth to be specific. More than half of that backlog is attributable to Anthropic and OpenAI, which, as I’ll say again and again, isn’t a problem because the money is yet to stop coming in.
Eyeballing the chart, something like $1.1T of the revenue backlog for Microsoft, Oracle, Google and Amazon is revenue that they expect to get from two companies that have yet to make a profit. That is a concentration risk.
Per data from fintech firm Ramp, 80% of OpenAI and Anthropic's enterprise revenues come from 1% of their customers, a number that hasn’t improved over the last three years. Ramp’s lead economist Ara Kharazian notes that the top 1% skews heavily toward the tech sector and AI products and services, and that this was a level of concentration risk unseen in any other software category they tracked.
benchmarking company Vals AI has been evaluating different open and closed models by using its own neutral harness. When every model ran on the same harness in the Terminal-Bench 2.1 evaluation, the open-weights model GLM 5.2 from Chinese company Z.ai (Zhipu AI) scored within a point of Anthropic’s Claude Opus 4.7 and 4.8 while costing about five times less per completed task.
In other words, paying for closed frontier models buys about a four-month head start at about five times the per-task cost—but only when currently looking at tasks taking [a human] between eight and 12 hours.
Below 8 hours there is no advantage for closed-source, above 12 hours neither model can reliably succeed. This isn't a great basis for OpenAI and Anthropic to project hundreds of billions of dollars of future revenue.
On the low end, that means that Anthropic and OpenAI account for over $200 billion dollars worth of expected revenues for Microsoft, Google and Amazon in 2027, which is contingent on their ability to raise venture capital or debt, which is contingent on the continued growth of their businesses, which is contingent on growing AI spend from a small subset of customers, many of whom are funded by venture capital.
The reason this hasn’t been a problem yet is that when you sign these contracts, you tend to pay a small up front fee, and the capacity in question is yet to come online.
...
In other words, Anthropic and OpenAI are currently in the teaser rate period where all of that capacity — and all of the associated costs — are yet to hit.
Next year, at least $200 billion in compute costs are coming due.
The question is whether Anthropic and OpenAI, two unprofitable, unsustainable AI labs that lose tens of billions of dollars a year, will be able to afford to pay them.
If they can't pay, about half of the revenue backlog for Microsoft, Oracle, Google and Amazon goes away. That reduces their EBITDA and potentially triggers downgrades.
Nangle starts with this chart, comparing the yields on the hyperscalers bonds with the average for similarly rated companies. It shows that:
it looks like the market is pricing Oracle and SpaceX as quality junk, and maybe Meta in line with BBB credit, and the rest of them as sort-of-in-line with the average AAA-AA-rated US corporate issuer. Is this cautious enough?
On recognized debt, the market’s calm is justified . . . [but] . . . [t]he real risk lurks below the waterline: off-balance-sheet debt lifts the debt burden by nearly 150% on average and pulls the model-implied credit quality down 1-2 notches.
There follows a complex description of Allianz' model that maps from stock prices to bond ratings, summarized thus:
Allianz calculates the ‘distance to default’ for the hyperscalers — a measure of credit riskiness calculated using equity inputs and balance sheet data — over time. And following work that Moody’s KMV has done mapping distance-to-default to default frequencies, they map the kind of ratings associated with these outputs.
Working solely off recognised balance sheet data, Allianz finds that the ratings implied by equity punters’ collective histrionics are higher than ratings assigned by the agencies for Alphabet, Amazon, Microsoft and Nvidia, a touch lower for Meta and a disaster for Oracle and SpaceX.
I think that's not quite right. Moody's gives Microsoft their highest rating, Aaa, and so does Allianz' model. The model gives Alphabet, Amazon and Nvidia Aaa, which is higher than Aa2, A1 and Aa1 respectively. But overall equity investors are (naturally) more optimistic than bond investors.
But these ratings are a testimony to the effetivness of the hyperscalers' off-balance-sheet techniques for distracting both the equity markets and the rating agencies. Nangle writes:
Once they whisk debtlike uncommenced leases into the equation, they find that Microsoft and Amazon drop into mid-investment-grade territory and that Meta falls into the quality end of junk.
Nangle's first chart showed the market already rating Meta at "the quality end of junk". But once the bond market assimilates the off-balance-sheet debts the downgrades and interest rate increases will be quite dramatic.
As stated in the paper, "A self-correcting multi-agent LLM framework for language-based physics simulation and explanation" [1], physics simulations are essential in science and engineering, but creating them often requires expert level knowledge of a variety of domains. For example, the development of a fluid simulation requires a strong understanding of the Navier-Stokes equations and the appropriate numerical solvers surrounding that area. In addition, users require proficiency in a programming language and familiarity with physics libraries, as building a simulation entirely from scratch is rarely, if ever, practical. The figure below illustrates the outcome of their research. Most notably, the user provides a basic ‘layman’ input and a basic set of parameters, while being afforded the luxury of omitting critical information required to develop a comprehensive solution. This is all accomplished using the Memory-Coordinated Physics-Aware Simulation (MCP-SIM), a self-correcting multi-agent framework producing a human readable output.
Figure 1: MCP-SIM Output
Traditionally, a “single-shot” has the anatomy of a user prompt, system prompt, and background information. Each of these components are combined to make up the context in which the Large Language Model (LLM) begins working on its output.
User Prompt: Initial input, request, and or instruction provided by a human user to the system. In this case, the type of physics simulation the user is looking to have created.
System Prompt: System instructions provided by the system’s designer helping define the context the LLM will use. It defines the persona, constraints, boundaries and ultimately the role before it processes any user input.
Background Information: Information that wouldn’t normally be stored within the LLM, but would be required to process the request.
Context: The collection of user prompts, conversation history, and additional environmental data produced by both the user and the model.
The primary limitation of this approach is the extensive information users must provide in their initial prompt. Furthermore, a single LLM is tasked with operating across multiple domains rather than maintaining a singular focus. This multidisciplinary scope introduces risks of cross-disciplinary terminology and variable confusion, further compounded by rapid context expansion that demands substantial VRAM to process.
To keep the system focused and to lower the barrier of entry for evaluating ideas, this paper introduces the MCP-SIM framework. This framework replaces simple one-shot requests that are sent to a singular LLM and instead builds a framework of multiple LLM based agents chained together in order to take advantage of an iterative cycle of expert evaluations. Each agent reviews the context, modifies, and enhances it to serve as input for the next. The cycle is defined by the figure below.
Figure 2: The MCP-SIM Iterative Cycle
MCP-SIM Core Architecture
The authors suggest an architecture that takes a user prompt and carries it through an assembly line of experts. Each expert adds or enhances based on their expertise to the context to produce a final output that is refined and constructed into a human readable format that’s delivered back to the user.
Figure 3: The MCP-SIM Workflow and Agent Collaboration
Specialized Agents
Input Clarifier Agent
The Input Clarifier is our Physics PhD Expert, and their job is to take a user’s input prompt and infer missing details such as domain geometry, boundary conditions, and initial conditions. A simple way to think of this agent’s purpose is that it generates a restatement of the layman’s description of the request and instead translates it to how a PhD’s view would define the request.
Code Builder Agent
The Code Builder Agent is an expert at writing Finite Element Computational Software (FEniCS) simulation code. Given the cleaned up prompt from the Input Clarifier, this agent then attempts to generate the appropriate simulation code.
Simulation Executor Agent
The Simulation Executor Agent’s job is to oversee the simulation. The Code Builder Agent delivered a set of blueprints for a FEniCS simulation and this agent’s job, for better or worse, is to see that they are executed to the letter. The code is executed and consequently evaluated for physical and numerical errors. At the end of this process, a quality assurance evaluation is performed and leads to one of two outcomes, on one hand, the simulation executes without critical runtime errors, syntax errors, or severe physical/numerical anomalies and the results are referred to the Mechanical Insight Agent. On the other hand, the simulation executes with a critical runtime error, syntax error, or a severe physical or numerical anomaly and is referred to the Error Diagnosis Agent.
Error Diagnosis Agent
The Error Diagnosis Agent is the clinician of the system. The Simulator Executor Agent detected an issue with the simulation, the executed code, collected the execution logs, tracebacks, numerical metrics, solver metrics, and any physical inconsistencies and referred them to be diagnosed. The Error Diagnosis Agent then determines whether it was an error in the code generation or an error in the PhD input. If the issue is determined to be an error in simulation code (most likely due to an error in the execution logs or the tracebacks), it is then resubmitted to the Code Builder Agent to attempt to correct the issue. However, if the error is with a simulation that fails to converge to a physical result, then the agent determines the issue is with its premise of the input and forwards the issue to the Input Rewriter Agent.
Input Rewriter Agent
The Input Rewrite Agent is the physics auditor of the framework. The Code Building expert did their job correctly and produced functional code, but the PhD overlooked a couple variables and possibly made a few mistakes. The auditor reviews the numerical and solver metrics, along with the physical inconsistencies and develops recommendations based on its findings. These findings are then forwarded to the Input Clarifier to rework the input into a more accurate description of the request and the process continues through the framework as before.
Mechanical Insight Agent
The Mechanical Agent is the multi-lingual publisher of MCP-SIM and its purpose is to generate a human readable report from the successfully run simulation. The report includes an explanation of the physics behind the simulation, what equations governed the simulation and a graphically generated output from the simulator.
Performance & Benchmarks
Benchmarks
The evaluation of the MCP-SIM framework was built around 12 benchmarks, where each level increases in complexity. The complexity of each level is tuned with the amount of prompt completeness and modeling complexity. These tasks are divided into 3 tiers: Simple, Intermediate, and Challenging.
Figure 4: The 12 Benchmark Prompts used [2]
Simple (Level 1-4)
The lower third of benchmarks are all designed to be fully specified, single physics problems. In other words, the user is providing a superior amount of information, in terms of the type of simulation that they want developed, along with exact variable values.
Intermediate (Level 5-9)
The middle third of benchmarks contains prompts that are incomplete and require additional inference for properties like geometry, materials, and/or boundary conditions. A great example is L6, where the prompt is looking to create an L-Shaped 2D pipe and to use that as a constraint to simulate viscous flow. This is more challenging for the LLM, since the viscosity of the fluid isn’t specified, not any of the geometry of the pipe included. The LLM must infer these missing parameters to successfully generate a simulation.
Challenging (Level 10-12)
Lastly, the most challenging benchmarks introduce multi-physics scenarios with missing assumptions based on published problems without existing code. Here, the LLM is expected to not only choose geometric and other physical properties, but also choose the correct set of physics models. For instance, the piezoelectric deformation under voltage is a mixture of material properties with electrodynamics.
Quantitative Performance
The performance of the MCP-SIM framework was measured against three other baselines. The first being, B1, a submission of the user prompt to a Generative Pre-Trained Transformer (GPT) as a single-shot request. The other two were implemented as stepping stones towards the full MCP-SIM framework. B2 being a single-shot enhanced with the clarifying input, and B3 extending it further into error diagnosis.
Figure 5: Number of Iterations of Success (L1-L12)
As seen in the figure above, the single-shot GPT setup (B1) consistently fails once the benchmark complexity reaches L7. With the exception of L4, the number of iterations required generally increases complexity.
Adding the Input Clarifier (B2) results in a relatively consistent requirement of four iterations until L6 where it begins to increase quadratically after discounting the Levels in which it was unable to generate a solution.
Following the inclusion of the Error Diagnosis Agent, the B3 baseline maintains a stable interaction count up to L10, again discounting for Levels in which it was unable to generate a solution.
Most notably, the Full MCP-SIM framework was able to develop a solution in each of the 12 benchmarks. While maintaining a low number of iterations up to L10, it experiences an increase rather rapidly for highly complex tasks. Nonetheless, the MCP-SIM framework demonstrates a clear increase in performance over the other three model baselines.
Comparisons
Examining purely the outcome of whether or not a baseline generates a solution, we can see that each sub-component of the MCP-SIM adds an increase in performance to each baseline. B1 baseline successfully generates a solution for the entire lower half of levels. The MCP-SIM framework significantly outperformed previous baselines, demonstrating that its self-corrective nature yields substantial performance gains.
Figure 6: Success Rate Across Baselines
Extending the Framework to Other Domains
Building upon the MCP-SIM framework, several extensions can be introduced to support a multi-domain expert ensemble capable of generating flight trajectories governed by physical dynamics and tactical mission objectives. This modular architecture enables each agent to maintain a concise and specific context. By restricting the context length, performance degradation in the language model is mitigated [3], thereby ensuring higher-quality outcomes. Another advantage of this design is parallelism. Assuming the input clarifiers are mutually independent, any number (N) of them can execute concurrently depending on hardware availability, rendering the increase in execution time negligible.
Although primarily a human-systems interface (HSI) study, this multi-domain, multi-agent implementation can be illustrated by modeling a system targeting objectives similar to those evaluated in the study "“Fly Like This”: Natural Language Interface for UAV Mission Planning"[2]. The synthesis of these papers is focused around generating two experts, the Aerospace Engineer and the Mission Engineer Agents. Both are designed to refine the user’s request into a capable set of mission parameters that ultimately develop a valid and effective flight path through the use of an MCP-SIM type framework seen below.
MCP-SIM Type Framework
Figure 7: Proposed MCP-SIM Framework Update
Specialized Agents
As before, each agent’s system prompt would be clearly and concisely defined allowing a user’s input to be placed into the assembly line of experts.
Aerospace Engineer (Input Clarifier): Translates high-level aircraft configurations into aerodynamic specifications (stall speeds, turning radii, lift/drag constraints). In addition, fuel or power usage rates of the aircraft would further constrain the full capability of the aircraft.
Mission Engineer (Input Clarifier): Translates high-level mission goals into optimized operational recommendations. To accomplish this, the agent would have to account for parameters such as terrain classification and geographical features, determining the most effective flight pattern for the mission.
Navigator (Path Generation): Given the limitations of the aircraft, this agent would generate a 3D Dubins flight path as defined by minimum turn radii, entrance and exit angles for each point. Furthermore, the output of this agent would be the evaluation of the combinations of various flight path segments allowing the agent to choose the most efficient path, in terms of distance and/or the most efficient flight path in terms of mission success.
Simulation Executor Agent: Tied into a simulation to flight the planned path by adjusting control surfaces of the aircraft. Locations at time (t) are tracked and logged
Error Diagnosis Agent: Reviews the logged data and determines how ‘on-course’ the drone was able to fly. If there are large divergences between the planned path and the executed path, it is examined and determined whether or not it was a fault in the Aerospace Engineer’s ability to determine flight characteristics or if the Navigator generated an impossible path to follow.
Flight Auditor: Upon a failed mission, this agent revises the input to be resubmitted to the set of input clarifying agents. This is accomplished by adjusting mission parameters or switching aircraft types (if possible). Could also recommend a change in the number of aircraft to be used or a change in altitude to perform the mission.
Record Generation (Mechanical Insight Agent): If the flight was successful, a report on the flight and its error (distance away from given path) as well as a graphical image of its journey is generated.
Conclusion
In this architecture, code isn’t necessarily generated, but a plausible flight path is determined based on aerodynamic features of the craft and what its designated mission is. A successful implementation of this modified framework could serve as either a basis to study a multi-domain, multi-agent MCP-SIM based framework or as a fourth user interface to be studied further building upon “Fly like this”. In either case, both avenues provide new areas of study on a path toward more robust and resilient interfaces.
We at OKFN are exploring how to connect AI to public data in a way that builds trust. This new publication shares our pilot process and invites you to test it in your own context.
In this role you will receive mentoring and community support to prepare for permanent appointments both inside and outside of academia. The successful candidate will work with Dr. Schneider to create an individual development plan.
The successful candidate will lead, conduct, and publish research, and write grant proposals, in an interdisciplinary environment. Key work will be to (1) conduct a scoping review of the literature; and (2) to develop and test scientometric indicators of epistemic integrity, methodological rigor, and innovation for the project CULTIVATE: Transforming Scientific Systems by Cultivating Virtue Ethics in Scientists. CULTIVATE is led by a team of three UW–Madison faculty members: Dr. Schneider (bibliometrics); a clinical psychologist at the Center for Healthy Minds (psychometrics); and an STS-inflected historian in the UW–Madison Department of Medical History and Bioethics (ethnographic and qualitative analysis to guide bibliometric development).
Key job responsibilities:
Lead, conduct, and publish research in consultation with Dr. Schneider
Design research projects, write grant proposals, and mentor junior researchers
Develop and execute a personal individual development plan to become eligible for permanent positions in your field, both inside and outside of academia, under the guidance of Dr. Schneider
The Information Quality Lab is a diverse group of researchers directed by Dr. Jodi Schneider. We value interdisciplinarity, sociotechnical perspectives, and diverse life experiences and world views. The Information Quality Lab invents tools and strategies for managing information overload in science, and for ensuring that high quality information is easy to find and use. The lab’s technical perspective draws on data science, argumentation theory, knowledge representation and reasoning, informatics, meta-science, computer supported cooperative work, and human-computer interaction. Technical skills commonly used in the lab include data science (network analysis, text mining, machine learning), knowledge representation (ontologies and semantic technologies), formal analysis and conceptual modeling, prototyping, annotation, mixed methods research, and user-based evaluations. Typical applications areas include digital libraries, health informatics, evidence synthesis, computer support for debate and argumentation, and bibliographic information retrieval (especially retrieval and quality of medical information).
We are part of The Information School (iSchool), which has been home to research and teaching that elevates questions of public good, people, and community in relation to computing, data and information for over 100 years. It is in a period of growth, more than doubling its faculty and enrollments over the past five years. It hosts the new Information Science undergraduate major as well as two professional master’s programs and a PhD degree. The iSchool is part of the new College of Computing & Artificial Intelligence which is committed to interdisciplinary and cross-campus research collaborations, establishment of innovative educational programs in the intersection of computing and other domains, and engagement with high-impact, real-world challenges.
Required Qualifications:
A PhD in any field (including, but not limited to, metascience, science of science, informatics, science policy, information sciences, information studies, library & information science, computational social sciences, etc.), earned within the last 5 years.
Interest in interdisciplinary research.
Excellent critical thinking, written and spoken English, and project management skills.
Preferred Qualifications:
Training and/or publications in fields relevant to meta-research or meta-science, such as scientometrics or data-intensive research about science.
Training and/or publications in formal literature reviewing methodologies.
Evidence of interdisciplinary research through scholarly publications or translational / implementation science projects.
Interest or experience in mixed methods research.
Interest or experience in grant writing.
Interest or experience mentoring undergraduate and graduate student research.
The ability to contribute to a positive, collaborative climate with students, faculty, and staff.
Compensation: $70,000 with benefits and paid holiday/leave for a 1-year full-time position. Anticipated start date: As soon as practicable.
To apply, submit your CV and cover letter with answers to short questions in this Google Form. Applications received by Wednesday September 23, 2026 are guaranteed consideration. While the position remains open, positions will be reviewed weekly each Monday.
EDIT September 29th, 2026: No further applications are being accepted.
LibraryThing is pleased to sit down this month with bestselling author Saratoga Schaefer, whose recent horror and thriller novels have caused quite a stir. Born in Brooklyn, they currently live in upstate New York, and made their authorial debut in 2025 with Serial Killer Support Group, described by author Dea Poirier as “a twisty and vicious thriller that will have you hungry for revenge.” 2026 publications include Trad Wife, The Last Time We Drowned, and A Thousand Monstrous Forms, a sapphic retelling of the horrific tale of Bluebeard in which a newly-married young woman must fight for her very survival. Schaefer sat down with Abigail this month to discuss A Thousand Monstrous Forms, published by Crooked Lane Books earlier in September.
Tell us about A Thousand Monstrous Forms. How did the idea first come to you? Did you always know it was going to be a retelling of Bluebeard?
A Thousand Monstrous Forms is a sapphic gothic horror Bluebeard retelling about a ceramic artist named Poppy Reed who moves into her new wife’s manor only to suspect it might be haunted. When she’s forbidden from entering the basement, Poppy finds herself drawn there over and over again.
This book was always going to be a Bluebeard retelling because I’ve always been strangely fascinated by the original folktale. But the impetus behind this book actually came many years ago when I took a medieval literature class in college. I read Edmund Spenser’s The Faerie Queene and fell in love with it, knowing I wanted to incorporate it into my own work one day. It took a long time to discover the bones of this book, but once I did, it came together in a way that really felt like kismet.
What is the significance of the Bluebeard story? Does recasting Bluebeard as a woman, and making the tale a sapphic one, change that story?
There are a lot of hints, nods, and easter eggs regarding the Bluebeard folktale in this book. It’s a retelling, so I wanted to be mindful of its source material while still creating a new story, especially one for the modern age. The theme of this book—realizing and recovering from abusive relationships—very much plays into the themes of the original Bluebeard. Making all the key characters women, and making the romance sapphic, was interesting because I got to explore new layers (such as queer partner abuse) not touched on in the original. Readers might be surprised to find how much that change affects the overall story…or not.
Can you elaborate on how you used Edmund Spenser’s The Faerie Queene in this book, and on any other influences in your story? Were there other classic tales that have shaped your book?
Spenser’s The Faerie Queene is a huge part of this book and brought all the themes and motifs together. I was researching Bluebeard and trying to figure out if there was a way I could also incorporate The Faerie Queene because I loved it so much. Imagine my surprise to discover that Spenser’s epic poem predates some of the iterations of the Bluebeard tale and might have even been the inspiration behind it. These two stories came together perfectly, allowing me to write my own. There are tons of references to The Faerie Queene within this book, including its title. Those who are familiar with the epic poem are going to see lots of winks and outright references and homages to the work within this book.
Your book speaks to issues of abuse and emotional trauma. What advantage does the horror genre offer to the storyteller, when exploring such themes?
Horror doesn’t shy away from darkness. In fact, horror is all about embracing the dark, analyzing it and compounding it until it’s even darker. Or until you feel some kind of catharsis from synthesizing your own pain in such a way. We sometimes think of horror as supernatural only—the things that go bump in the night or the monsters under the bed. But, for me, the true horror in all my books isn’t necessarily the spooky beings or jump-scares. It’s the humanity, the social aspect. How easy it is for us to hurt each other and how scarring that can be.
Tell us a little bit about your writing process. Do you have a particular schedule you keep, or rituals you observe? Do you know how your stories will end, when first beginning to write them?
I’m a morning writer for sure. I’m also a fast, intense drafter who needs to work in absolute silence. No background music for me—I get too distracted. Over many years I trained myself to write my first drafts in a little over a month because I learned that if I don’t get it done fast, I won’t get it done at all. I basically write two to three thousand words a day until the first draft is done, then I spend months (or sometimes years!) editing and revising. When I start to write, I always know how my stories will begin and end, but sometimes things in the middle change based on how the plot unfolds. Or if the characters decide to misbehave.
You’ve burst onto the scene recently, with four books in two years. What comes next for you? Do you have books in the offing that you can tell us about?
I didn’t initially plan on having three books out in one year, but since I’ve been writing novels for over a decade, I ended up with multiple manuscripts ready to go after my debut came out. I’m hoping to slow down a bit—eventually. I have two more announced projects I can currently talk about: my next psychological thriller, She’s a Big Fan, which comes out June 8, 2027, is about a fire-eater who notices strange behavior and disappearances at a pop star’s annual music festival; and Zeroth, my first sci-fi novel, out January 2028, about a starving artist who signs up for an android boyfriend beta test only to grow concerned by his dangerous malfunctions. Stay tuned…I’ll have some more news to share soon too!
Tell us about your library. What’s on your own shelves?
I’m a pretty diverse and eclectic reader; I like to read widely across genres. While of course I have a ton of thrillers and horror on my shelves, I also read sci-fi, fantasy, romance, literary fiction, and the occasional non-fiction. I’ve been on a cozy fantasy kick lately—fall makes me want to feel all warm and fuzzy inside.
What have you been reading lately, and what would you recommend to other readers?
It sure seems that a bunch of companies are trying to ship a git product of
some kind as of late. Wonder why that is.
Either way, I’m building a Git server backed by object storage as an
open-source project. It sounded simple
enough to start: Git looks like a filesystem, so let’s use a filesystem as a
translation layer on top of object storage to make Git speak object storage.
This model worked… ok, I guess? But
it didn’t work for real-world size repositories, so I needed a different
approach. Git stores everything in
Objects, so why not
store those as objects in Tigris?
Turns out Git packfiles and how they intersected with my (admittedly somewhat
terrible) filesystem shim were the main reason why it was slow. I ended up
having to invent my own packfile format with a columnar store that’s object
storage native. This is the fruit of all of my performance analysis, metrics
annotations, and more
Texas-style distributed systems work
than you can make your k8s cluster shake sticks at.
This approach worked surprisingly well for production-sized repositories, so I’m
sticking with this new Packfile format for now. It seems the least obtrusive
change to make Git Objects feel like object storage Objects, without any client
side changes.
When you make a commit, Git stores the changes you make as objects inside the
.git (I’ll call this “dotgit” so I don’t have to write as many backticks)
folder.
Imagine Git as two things: a sea of objects and named references to individual
objects. Each object is a
content-addressed
and compressed file. Here’s an example from a tiny git repository:
This produces several objects on the disk like this:
FIG 01A sea of objects, and a few names into it
.git/objects/ refs/heads/main
│
├── 1c/7a26a901..ec7966─────▶commit 1c7a26a
│ │
├── 8e/67afbb2e..857bd3─────▶tree 8e67afb
│ │ hello.txt
└── 9c/c9867337..09fe26─────▶blob 9cc9867
"Hello, blog!"
the filename is the sha1 of the bytes in the file, so the same
content is always, everywhere, the very same object
If you want to read the contents of an object, it’s compressed, so you have to
use a fairly evil looking python oneliner to scoop out the tasty innards:
As you can see, the objects are just bare files. Let’s look at a Git repository
of the Linux kernel and try to extract out an arbitrary commit. Everything
should just be a billionty bare object files, right? It should be easy to find a
single commit just by looking for the ID on the disk, right?
If only reality were so simple:
$ cd ~/Code/linux.git/
$ tree objects
objects
├── info
└── pack
├── pack-45986f41063f286029742ec12e2c2882b88c5786.idx
├── pack-45986f41063f286029742ec12e2c2882b88c5786.pack
└── pack-45986f41063f286029742ec12e2c2882b88c5786.rev
3 directories, 3 files
Yeah, as I’m sure you guessed just putting everything into their own files won’t
scale to something like the Linux kernel. I’m pretty sure you’d run into inode
limits like everyone did in the era of
fractal node_modules folders.
Aside
If you’ve used Node for long enough to remember that, please go get a
colonoscopy. Colon cancer is a real concern that too many people overlook for
too long and takes too many lives too early.
Git works around this by putting objects into
packfiles, compressed
bundles of objects that store them all in the same file. Here's an example of
the packfile efficiency in my checkout of objgit:
One of the beautiful things about implementing Git on top of object storage
like I am is that I’m using a platform where the object data is a sea of
objects with named references to points in that sea stored in FoundationDB.
This is a kind of divine recursion that I don’t really know how to describe
the beauty of. As above, so below.
Yo dawg, herd you like objects
Here's the object count for a copy of the Linux kernel:
This is eleven million objects, which at a very generous assumption of 10ms per
GetObject call means that fetching each of them takes over an hour to fetch them
all. The truth is there really aren’t 11M objects as individual files on the
disk, they’re bundled into one big happy 3.4Gi packfile. Your typical git repo
ends up accumulating them as it makes sense to break them up. My local copy of
the Tigris blog has 4 packfiles and 290-ish bare objects.
So you’d be thinking, “Oh, if git has packfiles, then why is the rest of this
post a thing?”
Well, like many things in distributed systems it’s complicated. Packfiles are
difficult because they’re designed with local storage and/or mmap in mind. Git
constantly writes packfiles to disk and then re-reads them. Filesystem reads in
that case are 10 nanoseconds at most (the filesystem cache helps so much here)
but doing any network roundtrip is 10 milliseconds at minimum. It’s at least a
million times slower because of how reality works.
Note
One of the things that /usr/bin/git does that makes integrating it into
object storage difficult is the unix-y idiom of writing to a file and then
immediately reading back from that file to calculate the hash. In object
storage you can’t GetObject something that hasn’t finished a PutObject call.
I worked around this previously by writing to the disk and then doing it that
way, but the experience kinda sucked in practice.
Messin' with Packfiles
Each packfile has an index that describes what’s in it. Here’s a view of the
index of the packfile made out of that trivial Git repo from earlier uppost:
$ git gc # force objects into a packfile
$ git verify-pack -v .git/objects/pack/pack-3971f5085c23c38be00e517ed0c64ca7df19b746.idx
1c7a26a901724b4ce766655ac387413fb9ec7966 commit 526 366 12
9cc9867337c2ebae85ba2350f901e0bcc209fe26 blob 13 22 378
8e67afbb2ee6bdcbb79061dfdfb93febce857bd3 tree 37 48 400
non delta: 3 objects
.git/objects/pack/pack-3971f5085c23c38be00e517ed0c64ca7df19b746.pack: ok
The commit points to the tree whose file “hello.txt” points to the blob and,
bob’s your uncle, you have a repo. Git uses these binary indices to let it know
where to look and how far it needs to seek into the packfile to know where to go
to get things.
FIG 02Eleven million objects, one packfile, one index
$ git count-objects -v.git/objects/pack/
count: 0└── pack-45986f41..c5786.pack 3.7 GiB
in-pack: 11827138
packs: 1eleven million loose files would be
size-pack: 3876775eleven million inodes. so: one file.
each row's colour is the run of bytes its offset points at, so a
read is a seek to an offset inside one very big file
ALL/03three rows, three offsets, three runs of bytes
The great part is that this works really well when everything is in a
filesystem. Git mmaps the packfiles so
that the kernel treats disk contents as memory pages, meaning that trying to
read past what’s “in memory” makes the kernel load it instead of userspace.
This is faster than loading it from the disk directly. It’s a shame this design
doesn’t work in object storage.
If only you could construct Range requests from packfiles
At some level this sounds pretty great for object storage, right? You have
offsets into the packfiles and then you can “just” grab out a single object from
a packfile with an
HTTP Range request
right? Objects are placed randomly within packfiles and other attempts at
storing Git in object storage end up having problems here. Tigris is really good
at random access scans, so most of the hard part is figuring out how to grab
the right data out of the bucket.
Aside
HTTP Range requests let a client download part of a file. The main usecase
they're built for is back in the day of dial-up internet you weren't online
all the time. Your main path to the Internet was the same way you send and
received phone calls. As such, if someone called you while you were online,
all your downloads got interrupted. Range requests let Internet Explorer
resume downloads where they got cut off instead of having to start all the
way over.
They've been maintained into the modern era but don't really get much use
outside of galaxy brain format abuse like what I'm doing and online video
streaming.
Well, it’s complicated. The example I gave shows all of the index entries one
after the other, but in the real world processing the index entries one after
the other you know the decompressed size of a single object in the packfile,
but not the compressed size. This means you don’t have enough information to
construct a HTTP Range request.
FIG 03One ranged GET, 366 bytes out of the middle of 128 MiB
This core problem is half the reason why I ended up needing to make my own
object-storage native Git packfile format.
SEND CUE SHEET
Way back in the days of physical media, one of the most common formats was the
CD-ROM (Compact Disc Read-Only-Memory, or CD). A CD is a 700Mi container that
stores data in sessions that each contain up to 99 tracks of either audio or
data. CDs were originally invented to store song audio in so that you could
listen to an hour or so of music at a higher quality than analogue cassette
tapes. CDs also let the player skip from track to track so you can go directly
to the song or movement of a larger work that you like.
Note
Of course this backfires if your CD mastering team decided to put
the entirety of Dancing Mad
into a single 17 minute track, meaning that you just have to know that
the blessed fourth movement is about 9
minutes into the battle music.
This is also why you see guides telling you to put legal backups of CD and DVD
media into lossless .iso files. An .iso file contains one recording session
that may contain data or audio.
FIG 04A CD cue sheet and an objgit cue sheet solve the same problem
│ TITLE "U0008_0000002_0000147_0002" │ ├──────────────────────────────────────────┤
└──────────────────────────────────────────┘ │ RECORD 1 58 B │
└──────────────────────────────────────────┘
to reach track 2 a player parses every
line above it. it already has the disc.record N sits at 16 + N*58. one seek,
then one ranged GET into objects.bin.
However there’s one catch that kinda ruins this easy way to back up CDs: they
can store multiple recording sessions on the same disc. Most of the time this
wasn’t used outside of making piracy on certain late 90’s/early 00’s game
consoles more annoying, but there was
that one Ricoh Encryptease product that combined
a user-recordable area with a factory printed area so that you could encrypt
files on CDs you share with the decryption software shipping alongside it. This
is about as cursed as it sounds.
Note
The trick of using multiple sessions is how Dreamcast games play as audio CDs
telling you to put it into a Dreamcast or how Xbox 360 games play as DVDs
telling you to put it into an Xbox 360. As an added bonus it means that when
you stick it into a computer it thinks that it’s an audio CD or DVD, which
means that lazy pirates can’t easily scoop out all the game files to their
hard drives.
As a result, there needed to be a way to properly handle this for archival
purposes. The eventual result was creating
cue sheets to store
alongside the binary blob of data. The .cue sheet stores information that the
decoder uses to be able to seek to arbitrary points in the .bin file. This
lets you easily extract things like songs or bits of data without having to read
the entire CD image. As an added bonus it handles multiple recording sessions
for you.
Packfiles v2: object storage boogaloo
This got me thinking, how would we take all of these lessons into heart and
build a new git packfile format optimized for object storage?
Wait, I know what you’re thinking. You’re thinking that I’m about to make a
Chesterton’s Fence violation.
Just “rolling my own” format for something as dear and precious as storing the
revision history of a company’s code repositories is probably one of the worst
decisions you can make, right?
Normally, yes, it’s a bad idea to do this. However Git is a distributed
version control system. When you clone a repository, you clone all of the
changes ever made to it on every branch at the same time. This also means that
everyone has a copy of the entire history of that repository, meaning that if
the worst does in fact come to pass and my handrolled format ends up sucking
it’s trivial to recreate all the data. Just push it again.
So what would this format look like?
Well for one the format needs to be Range-request native. You should be able to
scoop any one object out of a packfile without having to download or process
anything but the object you want. Again, Tigris is good at this, so we should
design the format with that usecase directly in mind. The format should also
take advantage of modern compression libraries like
zstd which are faster and more
data-efficient than zlib. Finally delta objects should be stored as their own
object in the packfile instead of slapped onto the end of the object it’s a
delta of so that you don’t have to read the object and its deltas to read the
object in the first place.
Note
I skipped over this earlier to save time, but Git stores both file revisions
(the entire copy of a file at any given point in time) and the difference
between them as an optimization to make it easier to uncompute the changes
made in commits. At some level this meme is both accurate and wrong:
Objgit's packfile format that probably needs a name
The format I came up with is pretty directly inspired from the .bin and .cue
format of CD backup. Objects are stored one after the other in a .bin file
that’s normally up to 128Mi (the oddly specific number was chosen because it
looked round to me) and the metadata of what objects are in there are stored
separately in a binary-encoded .cue sheet. Together this makes Git repository
storage in Tigris a columnar store.
FIG 05objects.cue on the wire: one header, then fixed-width records
objects.cue HEADER 16 bytes, once at the top of the file
The big thing I did was store the sizes for both the compressed and
uncompressed forms of objects alongside the offset into the packfile. This
means that you can trivially construct the right HTTP Range requests to scoop
individual objects out of Tigris while the packfile is downloading in the
background.
As an optimization for latency, whenever the Git library requests any object
from a packfile, the entire packfile is downloaded from object storage to a
temporary folder. Anything on the “far end” of the packfile is Range-requested
from Tigris until the downloaded packfile “catches up”. In practice this ends up
meaning that packfiles get downloaded just in time for them to be useful to read
from and there’s overall fairly little latency beyond what’s unavoidable with
the Git library I’m using. If the “scooping objects out of the far end” problem
ends up being an issue in practice I’ll just have it start fetching the most
recent packfiles for a given repository in the background when you start pushing
or pulling.
FIG 06Racing the beam: four ranged GETs while the container is 2 MiB in
four ranged GETs race the background download of the whole
file. whatever the download reaches stops needing a GET.
ALL/07the whole file lands, so B and D are cancelled mid-flight
Note
There’s nothing in the definition of the packfile format that limits packfiles
to 128Mi, I’m just doing that to prevent them from getting too big to download
quickly. In theory if you store a large binary blob (such as 3d models,
perfectly legal backups, etc) into the git repository directly it could result
in a packfile that’s bigger than 128Mi. I plan to solve this by implementing
Git Large File Storage in the near future.
But for now if you have a workflow that involves storing large binary files in
your git repository and want to use objgit for that: consider a different
architecture.
The funny numbers
One of the more surprising things about Git is that basically every interaction
with it is expensive. As such, you can go a long way by benchmarking how long it
takes to push/pull repositories. As such, I decided to compare against a few git
repos that have some interesting properties:
The big thing I wanted to fix was reducing the number of object storage calls.
Less object storage calls, less latency waiting for them to resolve. It ended up
being ridiculously effective. Object storage call counts sank like a stone.
Note
All charts in this post are at logarithmic scales so they render more cleanly
and the green bars are visible.
FIG 07S3 requests to push, before and after the format change (log scale)
1101001,00010,000
├─────────┼─────────┼─────────┼─────────┤
objgit
old████████████████████████231
new█████████████18
Xe/x
old████████████████████████████████████████9,236
new███████████████30
tigris-blog
old███████████████████████████████████3,324
new█████████████████████136
██oldgit packfiles behind a filesystem shim
██new.bin/.cue columnar packfiles
Repo
Build
Wall
S3 requests
PUT
GET
HEAD
LIST
Keys
Bucket bytes
objgit
old
8.7s (8.7s-15.9s)
231
46
47
0
138
42
829.27 KiB
objgit
new
2.2s (1.9s-2.5s)
18 (17-20)
5
10 (9-12)
0
3
4
902.98 KiB
x
old
3m29.4s (3m29.4s-3m40.6s)
9,236 (9,236-9,304)
1,087
2,170
0
5,979 (5,979-6,047)
1,082
54.96 MiB
x
new
14.3s (14.3s-14.7s)
30
6
20
0
4
4
46.38 MiB
tigris-blog
old
2m13.4s (1m29.1s-2m30.5s)
3,324 (2,687-3,417)
515
522
0
2,287 (1,650-2,380)
511
354.67 MiB
tigris-blog
new
26.5s (19.2s-27.4s)
136 (113-146)
9
123 (100-133)
0
4
8
360.75 MiB
The biggest gain was wall clock time for pushing though:
FIG 08Wall time to push, log scale, speedup on the right
1s10s100s1,000s
├───────────┼───────────┼───────────┤
objgit
old███████████8.7s
new████2.2s4.0x
Xe/x
old████████████████████████████3m29.4s
new██████████████14.3s14.6x
tigris-blog
old██████████████████████████2m13.4s
new█████████████████26.5s5.0x
██oldgit packfiles behind a filesystem shim
██new.bin/.cue columnar packfiles
One of the biggest places that objgit used to lag was pushing taking way longer
than it felt like it should. Eliminating the Tigris round trips made pushing way
more responsive.
Clone tests
I also wanted to see how the difference affected clone times. There were the
same benefits as with pushing:
FIG 09S3 requests to clone, before and after the format change (log scale)
1101001,00010,000
├─────────┼─────────┼─────────┼─────────┤
objgit
old█████████████████████████323
new████████████17
Xe/x
old██████████████████████████████████████6,428
new████████████17
tigris-blog
old████████████████████████████████████3,675
new██████████████████████158
██oldgit packfiles behind a filesystem shim
██new.bin/.cue columnar packfiles
Repo
Build
Wall
S3 requests
GET
HEAD
LIST
Wire bytes
objgit
old
11.8s (11.5s-19.5s)
323
51
0
272
763.63 KiB
objgit
new
2.6s (1.5s-5.7s)
17 (16-17)
13 (12-13)
0
4
767.96 KiB
x
old
3m23.5s (3m23.5s-3m33.6s)
6,428 (6,428-6,780)
1,091
0
5,337 (5,337-5,689)
40.68 MiB (40.68 MiB-40.70 MiB)
x
new
54.4s (54.4s-57.8s)
17 (17-71)
13 (13-67)
0
4
42.22 MiB (42.22 MiB-42.36 MiB)
tigris-blog
old
2m23.6s (2m19.4s-2m50.2s)
3,675 (3,601-4,123)
520
0
3,155 (3,081-3,603)
350.94 MiB (350.93 MiB-351.07 MiB)
tigris-blog
new
1m22s (1m20.8s-1m38.7s)
158 (137-317)
155 (134-314)
0
3
349.04 MiB (348.00 MiB-352.67 MiB)
FIG 10Wall time to clone, log scale, speedup on the right
1s10s100s1,000s
├───────────┼───────────┼───────────┤
objgit
old█████████████11.8s
new█████2.6s4.5x
Xe/x
old████████████████████████████3m23.5s
new█████████████████████54.4s3.7x
tigris-blog
old██████████████████████████2m23.6s
new███████████████████████1m22.0s1.8x
██oldgit packfiles behind a filesystem shim
██new.bin/.cue columnar packfiles
Oh yeah, the time numbers would probably be better if I tested this on a machine
with ethernet. I did all my testing with my corp laptop on Wi-Fi to specifically
put this in one of the worst conditions it could possibly be in.
Conclusion section
I’m still actively working on this. I’m not confident enough to use this for my
own projects yet and I wouldn’t blame you for not wanting to use this yet
either. I still haven’t implemented authentication, authorization, any kind of
API (my long-form SigV4 auth post was
actually going to be an objgit post!), or any rate limit beyond what your
machine can physically process. If you were to take this, run it, and then
expose it to the Internet, then anyone that can connect to that server can pull
or push whatever they want. Consider not doing that.
Objgit packfiles also currently accumulate forever, so if you have a bunch of
small pushes then there will be a bunch of small packfiles in the bucket. I’m
toying with designs that would occasionally compact them into bigger packfiles,
but that’s something that can be done later.
At the least though: Git’s packfile format is a great format for the constraints
of storing git repositories in actual filesystems. The moment you put network
roundtrips into the mix it all goes south.
I'm gonna keep working on this and publish reports like this as I learn more. I
hope this was interesting! Stay safe out there.
Open data leaders in Brazil and Uruguay jointly evaluate the pilot project they carried out with Open Knowledge, which ensures that reliable AI responses can be traced back to the original public data
Estimates are that, to justify the AI platforms' enormous capex plans, by 2030 they need to be generating around $2T/year in revenue. If every adult resident of the US spent $20/month on AI, it would generate $68.5B/year. Clearly, only the enterprise market stands even a remote possibility of generating the bulk of the $2T.
There are three major threats to the prospect of AI platforms extracting 6% of current US GDP from the enterprise market, and thus to OpenAI's and Anthropic's ambitions to IPO in the near future.
First, faced with AI's Affordability Crisis, companies have been placing strict limits on employees' spending on AI tokens.
Second, the gap in performance between expensive, closed-weight US models, such as OpenAI's and Anthropic's, and much cheaper, open-weight Chinese models has been rapidly closing, with the result that the US models are losing enterprise market share. Luz Ding, Spe Chen and Hayley Warren analyze this in US Lead in the AI Race With China Is Rapidly Narrowing:
Bloomberg in partnership with researchers at Vals AI, an independent AI evaluation and benchmarking platform, tested seven models from frontier Chinese and US companies to see how they performed in a real-world task. They were asked to create a fictional coffee e-commerce site called Brewberg using the same prompts. Most of the models scored 100% functional accuracy despite occasional design misses, but with very different price tags. The experiment employed the top performing models in July from Anthropic and all the Chinese firms, as well as more affordable models from OpenAI and Google.
They all did reasonably well, but the two best were Claude Fable 5 at $48.99 and Kimi K3 at $11.99. Chinese models charging much less for almost the same performance are grabbing market share:
the use of Chinese models overtook US platforms globally for the first time in June, and accounted for more than 60% of market share last month, on OpenRouter, a tech platform that offers software developers access to hundreds of AI models. It is a widely watched gauge of model usage despite tracking just a fraction of global AI consumption. The US, parts of Europe and Asia now favor Chinese labs, according to the same data.
On Hugging Face, Chinese AI models account for 41.4% of generative model downloads among developers, 5 percentage points higher than US models.
Third, it isn't just that the Chinese models are cheaper to run remotely, but also that because they are open-weight they can be run on affordable in-house systems, which means that:
They don't give Donald Trump a kill-switch for your busines.
They don't give Sam Altman or Dario Amodei a kill-switch for your busines.
They don't require giving the Chinese, Sam Altman or Dario Amodei all your business' critical data.
They provide visibility into and control over AI costs.
They are even cheaper.
The question is "compared to the closed-weight US models, what do you lose by running open-weight models in-house?" Below the fold I discuss a major study from Stanford that answers the question.
The 37-page paper is Intelligence per Watt: Measuring Intelligence Efficiency of Local AI by Jon Saad-Falcon et 14 al. Their abstract reads:
Large language model (LLM) queries are predominantly processed by frontier models in centralized cloud infrastructure. Demand growth strains this paradigm faster than providers can scale. Two advances create an opportunity to rethink it: small, local LMs (<=20B active parameters) now achieve competitive performance to frontier models on many tasks, and local accelerators (e.g., Apple M4 Max) can host these models at interactive latencies. This raises the question: can local inference viably redistribute demand from centralized infrastructure? This requires measuring both whether local LMs can accurately answer real-world queries and whether they can do so efficiently on power-constrained devices (e.g., laptops). We propose intelligence per watt (IPW), task accuracy per unit of power, as a unified metric for the capability and efficiency of local inference across model-accelerator configurations. We evaluate 20+ state-of-the-art local LMs, 8 hardware accelerators (local and cloud), and 1M real-world single-turn chat and reasoning queries. For each query, we measure accuracy (local LM win rate against frontier models), energy, latency, and power. We find three key results. First, local LMs successfully answer 88.7% of these queries, with accuracy varying by domain. Second, longitudinal analysis from 2023-2025 shows IPW improved 5.3x, driven by both algorithmic and accelerator advances, with locally-serviceable query coverage rising from 23.2% to 71.3%. Third, local accelerators achieve at least 1.4x lower IPW than cloud accelerators running identical models, revealing significant headroom for local accelerator optimization. These findings demonstrate that local inference can meaningfully redistribute demand from centralized infrastructure for a substantial subset of queries, with IPW serving as the critical metric for tracking this transition.
First, they ran a series of SLMs (QWEN 3, GEMMA 3, GPT-OSS, GRANITE 4.0) that can be downloaded on a local PC and compared their performance with cloud-based state-of-the-art LLMs (ChatGPT 5, Claude Sonnet 4.5, Gemini 2.5 Pro).
They ran these SLMs on local PCs powered either by an Nvidia chip or an Apple M4 chip, as they are readily available in current high-end desktop computers ...
Then they traced the performance of these SLMs vs LLM between 2023 and October 2025 on both chat tasks and reasoning tasks.
The results on four benchmarks are in Saad-Falcon's Figure 2, whose caption is:
Local Models Rival Cloud Models Across Diverse Benchmarks: Individual model performance scales with size, ranging from 31.5–69.4% for IBM GRANITE 4-H-S MALL, 30.0–83.6% for GEMMA 3-12B, 51.5–80.4% for GPT-OSS-120B, and 66.5–89.5% for GEMINI 2.5 PRO . Local routing (best local LM per query) achieves 97.8%, 88.3%, 77.0%, and 92.4% on WILDCHAT, NATURAL REASONING, SUPER GPQA, and MMLUPRO respectively, sur- passing cloud routing (100%, 82.9%, 66.5%, 87.4%) on three of four benchmarks.
which still make up the vast majority of requests today. As you can see, in every domain, the best SLM is able to find the same or better answers than an LLM in 90% or more of the cases, with an average across all domains of 98.6%.
It is really hard to justify spending 4-6 times as much for a 1.4% improvement in performance, so LLMs are no longer really necesssary for chat tasks.
which are obviously more demanding, SLMs are catching up fast. On average, they provide a better or at least as good an answer as LLMs in 62.5% of the cases.
SLMs may be catching up fast on reasoning tasks (see IBM's Granite 4.2 below) but it will be a while before LLMs are obsolete for these tasks.
the tasks for SLMs and LLMs are typically a mix of chat requests and reasoning tasks, so the third chart shows the weighted average of chat request performance and reasoning performance based on the frequency of tasks in each domain. As you can see, on average, SLMs are as good if not better than LLMs in 81.2% of the cases, with the LLMs having a significant advantage only in areas like engineering, life sciences, transportation and computer sciences.
But it’s not just accuracy. SLMs achieve this performance at energy and compute costs that are between 50% and 85% lower than for an LLM, depending on the SLM and hardware used in the computer.
The advantage for LLMs is being eroded quite quickly except for the extremely complex tasks. Saad-Falcon et al's Figure 7 shows how fast SLMs are catching up on reasoning tasks:
For reasoning tasks ... the pattern differs substantially. While levels 1-3 show strong improvements (+24.0, +37.8, and +53.9 pp respectively), levels 4 and 5 exhibit markedly slower progress. Level 4 improves by only +23.8 pp (7.93% to 31.72%), and level 5 remains largely unsolved with just +1.5 pp improvement (3.27% to 4.72%). This suggests that while local models have rapidly closed the gap on moderately difficult reasoning tasks, the hardest reasoning problems (those requiring either massive scale or capabilities beyond current architectures) remain a significant frontier. The presence of 134 level 5 problems (16.5% of the reasoning dataset) that remain 95% unsolved indicates substantial headroom for future model development in complex reasoning domains.
To understand why the rate at which SLMs catch up is critical we need to study Groundbreaker's The Teaser Period: Why the AI Boom Is Built to Break, which starts from the analogy between the subprime crisis and the AI bubble:
Paulson & Co. laid out the arithmetic that same month in a comment letter to the FDIC: Over 80% of recent subprime originations, it observed, were two- or three-year adjustable-rate products. The average subprime borrower’s mortgage payments already consumed roughly 40% of their gross income at the teaser rate. Almost none of them could service the reset rate out of income.
The crisis, in other words, was written in advance by the instruments themselves. The market looked at the reset wall and kept buying, because every participant believed the exit would arrive before the reset: home prices would keep appreciating and the borrower would refinance into a fresh teaser before the old one expired.
The take-or-pay compute contract - the instrument at the center of the AI build-out - has a structural feature that almost no one prices: its payments do not begin at signing. They begin at delivery. A lab signs a multi-year capacity commitment today, but the payments do not start until the data center is energized, the capacity is accepted, and the contractual ramp schedule commences - an interval set not by finance, but by construction: siting, powering, and filling a gigawatt-scale campus takes 24-to-36 months from signature - mirroring the two-to-three-year teaser of a subprime ARM.
More than $2.3 trillion of compute contracts now sit on the books of the four largest American cloud providers as remaining performance obligations and contracted backlog - signed, celebrated, capitalized into equity prices, and, critically, not yet billing.
...
The parallel to 2006 is exact and it explains the single most-cited absurdity of this cycle: How does OpenAI, a company with some $40 billion of run-rate revenue, sign $1.4 trillion of compute commitments? The same way a household with $60,000 of income signed a $600,000 mortgage: because the terms at signing do not require the payment yet, and because everyone at the table - borrower, lender, and the market - believes the growth will arrive before the payment does.
The chart shows that, in 2027 and 2028, the AI platforms will need to shell out $852B in cash for compute, whether they use it or not. We don't know how much revenue they are currently generating, but they want us to believe it is in the region of $100B. Ignoring all their other costs, they have to increase their revenue more than 4x next year to cover their contractual payments for compute. That means they have to extract at least $400B from the enterprise market in return for supplying it with technology that is slightly better than technology companies can run in-house around an order of magnitude cheaper.
What matters isn't the relative price/performance of in-house SLMs versus remote LLMs now, it is their relative price/performance when the LLMs' contractual compute payments come due, i.e. next year. It seems very unlikely that companies already balking at the cost of the AI platforms' products by moving to Chinese models will increase their spend 4x next year. It seems equally unlikely that investors will give the AI platforms a few hundred million dollars next year to burn so as to postpone the day of reckoning by another 12 months.
The Stanford study collected data in October 2025. Developments since, with more to come, have already significantly increased the price/performance advantage of SLMs. The include:
At Computex 2026 in Taipei on June 1st, CEO Jensen Huang announced the RTX Spark superchip — a single piece of silicon that combines a 20-core Arm CPU, a Blackwell GPU with 6,144 CUDA cores, and 128 gigabytes of unified memory, connected by NVIDIA’s NVLink chip-to-chip interconnect. The whole package delivers up to one petaflop of AI compute in a laptop form factor.
The number that matters: RTX Spark can run a 120-billion-parameter language model entirely locally, with a context window of one million tokens, without a single byte leaving your machine.
This is the guts of Nvidia's $5.2K DGX Spark desktop. It was Portent #20
Here’s a $94,011.50 desktop computer for sale right now.
It’s not a server rack or some crazy cloud infrastructure monstrosity.
It’s a (rather big) tower PC. Kinda like the one you played Cyberpunk 2077 on.
It sits under a desk and plugs into a wall like a regular desktop. The only difference is that it runs trillion-parameter AI models with no API keys, no per token payment, and no personal data leaving the room.
It is 18 times as expensive as the DGX Spark but can run models more than 8 times bigger. This was Portent 31.
Perplexity is launching Portable Computer today, a version of its agentic "Computer" platform that runs entirely on hardware users already own — starting with Nvidia's DGX Spark desktop supercomputer and Linux machines equipped with Nvidia RTX GPUs.
The launch, developed in close partnership with Nvidia, is one of the most aggressive attempts yet to move serious AI agent workloads off the cloud and onto local devices. The model, the user's files, and the work itself can all stay on the machine. Work completed locally consumes no billing credits, and the company says every task starts on the device by default — with the system asking permission before sending any individual step to a more powerful frontier model in the cloud.
Apple announced new iterations of both desktops, along with two new chips: the M6, the first 2nm chip in Apple’s M-series lineup for Macs, and the M5 Ultra, now the most powerful chip in the lineup for most things—especially AI workloads.
There aren’t any major new features for either machine. This is just a specs bump. But based on how Apple is presenting these refreshes, they’re leaning hard into those use cases, which weren’t even a thought when earlier iterations were first engineered.
The devices’ popularity for production inference took off after macOS 26.2 shipped last December. According to Apple’s release notes, 26.2 enabled “low-latency communication between Thunderbolt 5 hosts for use cases including distributed AI inference using MLX.” Thunderbolt 5 is a very fast wired data connection, and MLX is an open source array framework designed to help machine learning workflows take full advantage of the M-series chips’ unified memory.
Since then, both hobbyists and professional developers and researchers have been essentially daisy-chaining Mac minis or Mac Studios to run inference on local large language models that are much bigger than anything that could run a single mass-market device—providing an alternative to ultra-beefy specialized hardware featuring specialized Nvidia GPUs.
IBM has rolled out the newest models in its family of open-weight large language models designed to be downloaded and self-hosted. The newly launched Granite 4.2 comes in 3B, 8B, and 30B parameter variants.
Like previous versions, IBM is taking a decoder-only approach here. These new releases offer a 128,000-token context window natively. The 8B and 30B variants (not the 3B one) also go through an agentic reinforcement-learning block; they were trained for expanded capabilities like using the terminal, searching the web, or using external tools. The 3B model supports tools too, but without the same level of specialized training.
Beyond those tweaks, this release is particularly notable because, as IBM itself writes, “Granite 4.2 is the reasoning-focused release of the Granite language-model family.
Consider not just what Ding et al's graph of the narrowing performance gap between closed= and open-weight models, but also what Saad-Falcon et al's Figure 7 will look like after another year of both hardware and software developments like these.
Saad-Falcon et al argue that it isn't just the raw price/performance that advantages local compute:
System-level benefits offset per-query efficiency disadvantages. While cloud accelerators demonstrate 1.4× to 7.4× higher intelligence efficiency per query, local deployment provides complementary system-level benefits that offset this disadvantage. Local inference avoids datacenter infrastructure costs, network latency, and API pricing, while enabling 88.7% of queries that local models can handle correctly to bypass cloud compute entirely. As demon- strated in Section 4.3, intelligent routing between local and cloud infrastructure can achieve 60–80% reductions in total energy, compute, and cost compared to cloud-only deployment, even when local accelerators are individually less efficient. These findings suggest that the path to efficient AI infrastructure lies not in local accelerators matching cloud efficiency, but in routing systems that leverage the complementary strengths of both paradigms: local processing for the majority of straightforward queries and cloud infrastructure for the minority requiring frontier model capabilities.
It looks increasingly as though Klement is right that the hyperscalers are toast, because the vast majority of inference will happen locally while the massive data centers will be used only for training and for inference on massive models so expensive that almost no-one can afford them.
Update 4th September 2026
In the three days since I posted this, I have collected more evidence:
According to Business Insider, Thomson sits atop Snowdon, an intermediate model that Thomson Reuters built by reworking Qwen, an open-source offering from Chinese tech giant Alibaba. A joint team from Thomson Reuters and Imperial College London adapted Qwen over several months to ensure it was "ethically and politically de-biased and safe to use," Chief Technology Officer Joel Hron said.
Thomson Reuters spent roughly $40 million over two years on personnel and computing, the company said. The final training run cost approximately $450,000, according to SiliconAngle. The company chose to forgo developing a foundation model from the ground up, instead taking an existing open-weight model as its starting point and enriching it with proprietary content, specialized training methods, and domain expertise.
Cloud-based AI models operated by OpenAI, Anthropic, xAI, and Google suffered a rare and overlapping set of significant service interruptions over a period of hours Thursday morning.
Anthropic first reported a “partial outage” related to “elevated errors on requests to Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5” at 9:23 am (all times Eastern). The company reported that it had “identified the cause” of the error roughly 15 minutes later, before reporting that “a fix has been deployed” and the issue was resolved by 12:16 pm. A separate incident report indicated “elevated errors on requests to Claude Sonnet 5” for a brief period just after noon.
OpenAI, meanwhile, reported “elevated errors across ChatGPT and Codex” were resulting in “degraded performance” as of 10:43 am Thursday morning. A mitigation put in place a little more than half an hour later led to the issue being marked as “resolved” by 12:55 pm.
As of this writing, xAI’s Grok currently displays a user-facing error message that the model “is experiencing issues” and that the company is “working on restoring service as quickly as possible.” User-submitted reports from DownDetector regarding Grok shot up from less than 10 just before 9 am to 1,365 by 9:45 am and have fallen to 273 as of this writing.
While Google has not publicly acknowledged any issues with its Gemini model, reports of problems on DownDetector similarly spiked from just 23 around 10:30 am to 412 just after 11 am. Data collected by API checker StatusGator also shows what it terms a “likely outage” for the Gemini API between 10:45 and 11:15 a.m. before the status returned to normal.
Other Internet services did not experience problems, so this seems to have been solely an AI problem.
NVIDIA Personal AI Router (PAIR) is software that connects compatible macOS, Windows, and Linux systems with NVIDIA RTX™GPUs and DGX Spark systems into a personal home AI cluster. PAIR distributes local AI inference workloads across available devices while keeping prompts, files, and agent context on the user’s home network.
Hybrid Compute on Apple silicon orchestrates a task between frontier intelligence in the cloud and a local model on the Mac. Cloud models handle research and reasoning, while a local model works with private files and apps on the Mac.
For this division of labor to feel seamless, local inference must keep pace with the rest of the task. That requires an engine that can process prompts quickly and sustain a high token-generation rate.
Lily, our lightweight local inference engine, is built specifically for Apple silicon and Qwen3.6-35B-A3B, with separate optimizations for prefill and decode. A standalone demo is publicly available on GitHub.
Pair and Lily illustrate the rapidly growing trend of moving as much inference as possible to local hardware and, relatedly, the importance of placinng a router bewteen the user and the inference systems.
One of the ‘known unknown’ threats that S&P Global Ratings outlined last week hanging over the biggest companies driving the AI juggernaut is the notion that open-weight models close the performance gap with expensive frontier models.
We can break this down further into two component risks. First: that some queries put to frontier LLMs — which require centralised cloud infrastructure — can be answered more quickly and more cheaply by locally hosted set-ups, fitting on something maybe only a little larger than Balakrishnan’s Raspberry Pi. And second: that big complex open-weight models become pretty indistinguishable from big complex closed-weight models, can be accessed at a fraction of the cost, and come with riders that are valuable in their own right.
As regards the small local models, Nangle both conducts his own experiment and also relies on the Saad-Falcon et al paper. He concludes:
So we can see how SLMs might appeal to management seeking to repair their bottom lines and reputations for cost control after token-maxxing experiments backfired so spectacularly.
In fact, given quite how basic most AI queries appear to be, the researchers’ results sound like the sort of thing that might even challenge the economics of building some of the myriad mahoosive data centres being planned.
As regards the open-weight LLMs, he acknowledges that relative to the frontier closed-weight models their performance lag is shrinking:
Cloud-based open-weight models, just like their closed-weight counterparts, gobble up data centre processing capacity. As such, it’s hard to see what direct problem they might pose to the economics of data centre build-out. Although it’s easier to see how they might be a problem for AI labs like Anthropic and OpenAI.
...
Good luck though trying to download Moonshot’s Kimi K3 on to your PC. Operating across 2.8tn parameters, it’s just a different animal. According to Citi Research, Kimi K3 trounces every closed-weight frontier model ever built prior to [checks calendar] three months ago.
It is clear that almost as good but vastly cheaper is winning in the market:
One way to see which way the wind is blowing on open-model versus closed-model usage is by looking at data from router firms like OpenRouter, the New York start-up that Stripe agreed to buy last month.
...
Increasingly, the share of queries that are truly closed-weight frontier-model-worthy is declining. Back at the start of the year, three-fifths of its queries were routed through to closed-weight proprietary models. The latest share is just a quarter.
Among the companies Nangle cites as moving to open-weight models are DoorDash, Siemens, Airbnb, AT&T, Latham & Watkins and other major law firms.
As with Thomson-Reuters, many are concerned with security:
Tareq Islam, a strategic adviser to ApexE3, a capital markets AI infrastructure firm, tells Alphaville that some London-based asset managers, as well as not wanting to build dependency on any single proprietary model, are reluctant to post their most confidential data and intellectual property to closed-model providers. Clients of ApexE3 include Vanguard, the world’s second-largest asset manager.
OpenAI claimed to have found a “singularity” in the Navier-Stokes equations in three dimensions, one of the six unsolved “Millennium Prize Problems” set by the Clay Mathematics Institute. In a statement posted to X.com the lab wrote, “The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra.”
A few hours earlier, Tristan Buckmaster, a professor of mathematics at NYU who has been working on closely related problems for some time, posted his own statement in which he went to great lengths to NOT EXACTLY formally accuse OpenAI of stealing an almost complete proof of Navier-Stokes from his sessions on Codex, GPT-5.6 Sol, and Astra. But he did not sound at all happy, and it’s worth a read in full if you’ve got five minutes.
Agent memory has been touted as a dimension of growth for LLM-based
applications, enabling agents that can accumulate experience, adapt
across sessions, and move beyond single-shot question answering. The
current generation of agent memory systems treats memory as an external
layer that extracts salient snippets from conversations, stores them in
vector or graph-based stores, and retrieves top-k items into the prompt
of an otherwise stateless model. While these systems improve
personalization and context carry-over, they still blur the line between
evidence and inference, struggle to organize information over long
horizons, and offer limited support for agents that must explain their
reasoning. We present Hindsight, a memory architecture that treats agent
memory as a structured, first-class substrate for reasoning by
organizing it into four logical networks that distinguish world facts,
agent experiences, synthesized entity summaries, and evolving beliefs.
This framework supports three core operations – retain, recall, and
reflect – that govern how information is added, accessed, and updated.
Under this abstraction, a temporal, entity aware memory layer
incrementally turns conversational streams into a structured, queryable
memory bank, while a reflection layer reasons over this bank to produce
answers and to update information in a traceable way. On key
long-horizon conversational memory benchmarks like LongMemEval and
LoCoMo, Hindsight with an open-source 20B model lifts overall accuracy
from 39% to 83.6% over a full-context baseline with the same backbone
and outperforms full context GPT-4o. Scaling the backbone further pushes
Hindsight to 91.4% on LongMemEval and up to 89.61% on LoCoMo (vs. 75.78%
for the strongest prior open system), consistently outperforming
existing memory architectures on multi-session and open-domain
questions.
The Digital Preservation Coalition (DPC) has today released its complete
three-part Technology Watch Guidance Note series, Cyber Security and
Resilience for Digital Preservation, for general public access.
Following an exclusive preview for DPC Members over the summer, all
three Guidance Notes are now freely available to the wider digital
preservation community.
Turn any codebase, with its docs, SQL schemas, configs, and PDFs, into a
queryable knowledge graph. A /graphify skill for Claude Code, Cursor,
Codex, and Gemini CLI: local deterministic AST parsing, every edge
explained, no vector store.
Traditional RAG systems retrieve what was explicitly said, but they miss
what matters most—the insights only accessible by rigorously thinking
about your data. Without reasoning, you’re leaving latent information on
the table. Static retrieval can’t surface implicit connections,
struggles when new information contradicts old data, and fails when you
need to make predictions under uncertainty. Honcho uses formal logic to
extract all that latent information. This reasoning is AI-native—it
performs the rigorous, compute-intensive thinking that humans struggle
with, instantly and consistently. The result is memory that goes beyond
simple RAG recall to provide exhaustive context for statefulness.
Hindsight is an agent memory system that provides long-term memory for
AI agents using biomimetic data structures. Unlike traditional RAG
(Retrieval-Augmented Generation), Hindsight:
Stores structured facts instead of raw document chunks
Builds mental models that consolidate knowledge over time
Uses graph-based relationships between entities and concepts
Supports temporal reasoning with time-aware retrieval
Enables disposition-aware reflection for nuanced reasoning
The blueprint for modern digital computing was codesigned by Charles
Babbage, a vocal champion for the concerns of the emerging industrial
capitalist class who condemned organized workers and viewed democracy
and capitalism as incompatible. Histories of Babbage diverge sharply in
their emphasis. His influential theories on how “enterprising
capitalists” could best subjugate workers are well documented in
conventional labor scholarship. However, these are oddly absent from
many mainstream accounts of his foundational contributions to digital
computing, which he made with mathematician Ada Lovelace in the
nineteenth century.1 Reading these histories together, we find that
Babbage’s proto-Taylorist ideas on how to discipline workers are
inextricably connected to the calculating engines he spent his life
attempting to build.
It’s a dedicated Guix channel from the Guix-Science project. The channel
provides a comprehensive, community-driven, scientific software catalog
suitable for the global scientific community.
With the increased usage of Guix in scientific contexts, there is also a
growing need for packaging software used in the research and teaching
spheres. Although the Guix project encourages users to contribute
packages to the main Guix channel, some packages cannot be included
there because they do not comply with Guix’s packaging policy, or
because they are too specialized to be used outside an HPC context. The
Guix-Science channel has a more permissive policy than the main Guix
channel; especially regarding the usage of pre-built components.
Additionally, it has a more lenient deprecation policy than Guix proper.
What is a Civboot? Conceptually it is a seed from which can grow the
essential technology of Civilization. Imagine a factory capable of
building a minimalist general-purpose personal computer, like something
from the 80’s/90’s or better. A Civboot factory is one which can build
such a factory from raw materials.
The first step in the Zen of Reticulum is to realize that there is no
cloud. There is only other people’s computers. When you build for the
cloud, you are building for a landlord. You are accepting that your
application’s existence is conditional on the permission, uptime, and
continued goodwill of a central authority.
In Reticulum, you must shift your thinking from “connecting to” to
“being among”. Reticulum is not a service you subscribe to - it is a
fabric you inhabit. There is no “up there”. There is only here and
there, and the space between them is peer-to-peer.
This paper proposes decomputing as a programme for pushing back against
such technocratic nihilism. Decomputing adopts a degrowth framework and
advocates for reclaiming agency through deautomatisation and people’s
councils, in order to develop convivial alternatives that centre an
ethics of care. It is a philosophical and tactical response to the wider
crisis of care induced by the collapse of existing systems.
Decomputing proposes a prefigurative praxis for those experiencing
ruptures in the infrastructures of everyday life, while minimising the
opening for an authoritarian turn or war economy. Rather than
capitulating to an increasingly fascistic world order, decomputing
advocates a shift in the idea of order itself, one that means
challenging the idea that ever more advanced tech is synonymous with
progress.
This interview was conducted in writing with Autistici/Inventati
following its August 26, 2026 designation by the United States under
Executive Order 13224.
The conference aims to create a forum for sustained exchange between
those who preserve born-digital materials and those who interpret them
as historical sources. By bringing together archival, methodological,
technical, and historiographical perspectives, we seek to ask not only
how born-digital records can be saved, but also how they can be made
meaningful for contemporary history. We especially welcome contributions
that combine methodological reflection with concrete case studies,
project presentations, or hands-on approaches.
Large-language models (LLMs) are rapidly becoming part of human culture,
reshaping how information is produced, transmitted, and used. Here we
propose that their diffusion can be understood through a viral analogy,
with LLM use spreading through populations, becoming embedded in
cognitive and cultural practices. We model transitions among uncoupled,
coupled, and persistently dependent users, and show that the interplay
between social transmission, recovery, and collective reinforcement can
generate tipping points and technological lock-in. A central consequence
is the possibility of runaway dynamics: once a critical threshold is
crossed, small increases in adoption can trigger rapid population-level
shifts toward persistent dependence, with abrupt losses in cognitive
competence. The same framework, however, identifies conditions for
cognitive immunization, based on reducing transmission and facilitating
reversibility. Our results highlight how LLM adoption may involve
nonlinear collective transitions with important consequences for
cognitive autonomy.
You should fully expect your writing to be run through Pangram. If your
position is that we should be fine with an LLM crafting prose from your
prompt, spare us all the wasted cycles and just give us your prompt. Or,
better yet, consider doing what generations of writers have done before
you, and treating that prompt as a skeleton that you use to write your
piece yourself!
I’ve made a small website called Five Thank Yous.
It is a quiet, curated collection of anonymous gratitude threads: short reflections that begin with one specific good thing and trace it back through the people, labor, nature, place, memory, timing, chance, and care that made it possible.
(From the project's about page:) The idea started in my personal journal.
I'd been keeping a small morning gratitude practice — five bullet points each day — when one morning I found myself tracing one of them backward instead of moving to the next.
I think it was gratitude for clean drinking water.
Rather than writing five separate things I appreciated, I found myself asking what made the previous thing possible: the water, the treatment plant, the workers, the scientists, their educators.
A few weeks later I remembered Toyota's "Five Whys."
In hindsight, it seemed obvious to turn that same structure toward gratitude instead of problem-solving.
In the summer of 2026, I was looking for a project to learn some new skills and put something good out into the world.
Creating a Five Thank Yous website seemed like a manageable side project, and I wanted to gain experience building static websites using AWS technologies.
The basic idea is simple.
Start with one concrete gratitude. Not “I’m thankful for food,” but something more particular:
Thank you to the bowl of soup waiting on the stove after a cold walk home.
Then ask:
Who or what made that possible?
Maybe the next thank-you is to the person who cooked it.
Then to the recipe they used.
Then to the magazine or website that published the recipe.
Then to the writer who contributed the recipe.
A Five Thank Yous thread follows one line of connection:
Thank you to one specific good thing.
Thank you to what made that possible.
Thank you to where the thread leads.
It is not a list of unrelated things.
It is one thread.
Each thank-you follows from the one before it.
Why I made it
Gratitude is often treated as a private feeling: I feel thankful for something good in my life.
But gratitude can also be a kind of recognition.
It can help us notice that the good in our lives does not come from ourselves alone.
A cup of coffee may lead to a roaster, a farmer, rain, soil, insects, transport, ceramics, electricity, a quiet kitchen, or ten minutes of peace.
A repaired chair may lead to a neighbor, a tool, a YouTube video, a tree, a craft tradition, patience, or the decision not to throw something away.
A text from a friend may lead to years of friendship, a moment of courage, therapy, recovery, someone else’s kindness, or the luck of having met.
The point is not to arrive at the "correct" root.
There is, in fact, usually no single root.
The point is to make the connectedness visible.
Those are not missing features.
They are design choices.
I wanted this to be a calmer kind of public space: less like a feed, less like a performance platform, and more like a small collection people can visit when they want to slow down for a few minutes.
Submissions are anonymous, but they are not automatically published.
Everything goes into a private moderation queue first.
I review entries one-by-one and may lightly edit them for length, clarity, formatting, privacy, or safety before publishing.
That means publication is not guaranteed.
It also means the site is not an open anonymous wall.
The goal is to keep the collection careful, safe, and coherent.
Try it yourself
If you want to try the practice, here is the simplest version:
Start with one specific good thing.
Ask: who or what made that possible?
Make that answer the next thank-you.
Ask again.
Follow one thread until you have a thread of thank-yous. (Five is a rough guideline, not a goal or a mandate.)
If you write a Five Thank Yous thread and would like to offer it to the public collection, you can submit it anonymously.
Of course, I hope (and expect) the act of writing a thread will be worthwhile even if the entry is never published.
Please share it if it fits
I’m sharing this in mid-September because November is approaching, and in the United States that means many people will soon be talking about gratitude.
Thanksgiving is a natural moment for a project like this, but I do not want Five Thank Yous to be only a Thanksgiving site.
Thanksgiving can be joyful, complicated, lonely, conflicted, meaningful, painful, or all of those at once.
Gratitude itself can be complicated too.
My hope is to spend the next couple of months slowly sharing the site, gathering more submissions, and publishing new entries as Thanksgiving arrives.
If Five Thank Yous reminds you of someone — a writer, teacher, librarian, minister, therapist, chaplain, community organizer, group facilitator, or person who likes quiet corners of the internet — I’d be grateful if you passed it along.
The project is not trying to become a social platform.
It is not trying to maximize engagement.
It is not measuring success in pageviews.
The best outcome would be a small, thoughtful collection of sincere entries that people can read slowly.
Ten beautiful submissions would mean more to me than ten thousand drive-by visits.
The RecordDataFormatter is one of VuFind's most powerful components for controlling how it displays bibliographic metadata..
Over the years, it has evolved significantly, with much of its configuration moving out of PHP code and into an .ini configuration file.
These are resources and slides from a presentation at WOLFcon2026.
Thomas Wagener gave a fresh overview of the RecordDataFormatter's architecture and capabilities, including how to extend it to non-default record drivers such as EDS.
I introduced a practical "all-config" strategy that moves every code-defined field specification into INI files (e.g., RecordDataFormatter/DefaultRecord.ini).
That dramatically simplifies day-to-day management of display fields while preserving the ability to inject advanced logic—like custom multiFunction callbacks for authors and multi-line fields—through a thin PHP override layer.
Learning objectives from the session:
Understand the RecordDataFormatter pipeline — how spec keys, field definitions, helper methods, and rendering templates work together to turn record driver data into displayed metadata.
Apply the RecordDataFormatter to additional record drivers — learn techniques for reusing or adapting display specifications for non-default sources like EDS.
Implement an all-config strategy — move field specifications out of PHP and into INI files by overriding PHP to return empty arrays, letting the INI file be the single source of truth.
I shared the presentation with Thomas Wagener from HeBIS.
Thomas developed the RecordDataFormatter functionality in VuFind, and in his part of the presentation he showed how the functionality works and three ways to configure the display: through RecordDataFormatterFactory overrides, via Spec Plugins (introduced in VuFind 11), and INI configuration files.
My part of the presentation followed Thomas', and I showed how to extend the RecordDataFormatter so all fields are defined in INI files.
Slide 20: Option 3-alternative: Use an All-INI Configuration
As Thomas showed, while PHP Spec Plugins and INI configurations give us significant control over RecordDataFormatter, there’s an alternative pattern that pushes this flexibility even further.
Instead of scattering field definitions across PHP classes or using INI files merely as overrides, the All-INI approach exposes all spec options directly in configuration files.
There are two primary drivers here. As a hosting provider, much of our implementation process revolves around making VuFind match the needs of a library's collection, and that involves how fields are displayed.
For libraries with only one VuFind installation, the primary driver may be empowering site administrators and metadata librarians.
They can customize record views, reorder fields by adjusting the INI list sequence, or move fields between metadata contexts by cutting and pasting sections—all without editing or deploying PHP code.
To understand how this works, let’s look at the basic anatomy of a RecordDataFormatter configuration file.
Slide 21: Anatomy of a RecordDataFormatter Configuration File
A standard configuration file breaks down into five distinct sections:
[Global] sets site-wide baseline defaults such as default separators or global display toggles.
[Defaults_Function_Mapping] maps metadata contexts like core or description to default spec methods.
[Defaults] explicitly lists which fields belong to each section using arrays like core[] or description[].
[Field_<Name>] blocks define specific options for individual fields—such as dataMethod, renderType, template, itemPrefix, and itemSuffix.
[overrideContext] and extraLineOptions handle advanced conditional rendering rules.
Now, let's contrast this standard hybrid configuration against an all-INI strategy.
Slide 22: Standard Approach vs. All-INI Strategy
In the standard approach, PHP SpecBuilder classes like DefaultRecord.php remain the primary source of truth for metadata specs.
The placement of fields in the core context and description tab context are hardcoded in PHP methods, and field reordering requires calculating explicit pos values.
In the All-INI strategy, 100% of field definitions, context assignments, and render rules are declared in configuration files.
Instead of doing integer math for pos values, the sequential position is assigned based on the visual order of entries in the INI file.
The PHP class becomes a lightweight bridge for dynamic closures and other things that can't be specified in static INI configuration files.
Let's look at how this field-to-context assignment looks in practice.
Slide 23: Moving All Field Declarations into INI: Field-to-Context Assignment
Here under [Defaults], we explicitly list every field for the core metadata display and the description tab.
In a couple of slides, you’ll see that we are removing the field definitions from the PHP methods; everything rendered on the interface is declared here in configuration.
Notice how Published in, Authors, Language, and Edition are grouped under core[].
At the same time, Summary, Item Description, and Access reside under description[].
Once the fields are assigned to contexts, we configure their formatting rules.
Slide 24: Moving All Field Declarations into INI: Field Definitions
Each field is given its own [Field_*] configuration section.
If you've defined new fields using the INI file, this likely looks familiar.
With the All-INI approach, we define all fields in the field configuration section.
For a basic field like Access, we simply declare dataMethod = "getAccessRestrictions".
For complex fields like Published in, we specify dataMethod = "getContainerTitle", renderType = "RecordDriverTemplate", and point directly to the template data-containerTitle.phtml.
We can also handle Microdata inline, as shown in [Field_Edition], where we wrap values with Microdata spans using itemPrefix and itemSuffix.
This explicit structure dramatically simplifies day-to-day administrative tasks.
Slide 25: Simplifying Reordering & Tab Transfers
Consider how much simpler routine maintenance becomes:
To reorder fields visually, you move lines up or down in the INI file—the underlying PHP bridge automatically re-sequences pos values in increments of 100.
To move a field like Item Description or Access from the core section to the description tab, you don't need to extend PHP spec methods—you move the field string from the core[] list to the description[] list in INI.
This achieves true separation of concerns between code logic and display presentation.
To make this work, we use a custom PHP class that acts as a lightweight bridge.
To get the all-INI option to work, we do have to do some PHP code.
All code shown on this and the following slides is published in at the Gist in the resources section above.
To hand total authority over to the INI configuration files, our custom DefaultRecord.php class overrides getDefaultCoreSpecs() and getDefaultDescriptionSpecs() to return empty arrays.
By clearing out the hardcoded PHP specs, VuFind relies entirely on what is declared in configuration.
Next, let's see how this bridge class automatically assigns positions and attaches dynamic callbacks.
In getDefaults(), we iterate through the field specifications parsed from INI.
First, we automatically assign sequential $pos values starting at 100, incrementing by 100 for each field to respect the exact visual order in the INI file.
Second, for complex fields requiring dynamic PHP logic—like Authors, Language, or helper methods for Format_CORE—we attach the corresponding PHP closure callbacks.
We also include fallback handling for multi-line fields and data type normalization.
If a field specifies renderType = 'Multi' in the INI without a custom multiFunction parameter, our bridge automatically attaches a default getLabelDataMapFunction() callback.
We also handle INI-to-PHP string normalization: if dataMethod is parsed as the string "true", we cast it to a boolean true so template-only renders function correctly.
Let's examine that multi-line label-data processing function.
Slide 29: Callback Function for Processing Label-Data Pairs
The getLabelDataMapFunction() method returns a callable closure designed for multi-line fields.
It iterates over raw record driver data arrays, extracts label and value pairs, and sets sub-positioning options ($options['pos'] + $i) for each rendered line.
This ensures multi-line fields render with predictable sub-sorting without cluttering configuration syntax.
Finally, there is one small translation detail to handle when defining context-specific fields.
Slide 30: Translation File Additions
It wasn't until I got into this work that I realized that both the core context and the description tab context have fields labeled "Format" and "Published".
This is a problem when we create our field definitions in the INI file — how does the RecordDataFormatter class know when we are referring to the field definition that is traditionally in the "Core" context versus the one that is traditionally in the "Description" tab context.
We do that by putting suffixes on the field names when we list them in the defaults and when we define them in the field definition section.
That solves our disambiguation problem in the configuration, but we also want both to display standard, user-friendly labels on the screen.
In languages/en.ini, we map both contextual keys to their clean display labels: Published_CORE = Published and Published_DESCRIPTION = Published.
This preserves clean, distinct keys in configuration while maintaining uniform presentation for patrons.
Let's summarize everything we've covered today and look at where this model can go.
Slide 31: In Summary...
To recap: RecordDataFormatter provides a powerful engine for transforming record driver getters into structured metadata displays across contexts like core, description, result lists, and EDS.
Thomas showed how to customize the display with factory overrides, Spec plugins, or hybrid INI configurations.
The All-INI strategy expands configuration to 100% of field lists, context assignments, and ordering.
Looking ahead, a future direction for this work could be dynamic MARC field extraction.
By developing a dynamic dataMethod syntax directly in configuration—such as dataMethod = "marc:245:a:b"—we could eliminate the need to write custom MARC getter methods in PHP modules entirely.
VuFind's local customization model encourages copying core configuration and theme files into a local directory, keeping institutional changes separate from the upstream codebase.
While this simplifies understanding local modifications, it creates a significant maintenance burden during upgrades: local copies drift out of sync with evolving core files, and reconciling changes by hand is tedious and error-prone.
These are resources and slides from a presentation at WOLFcon2026 on how to use kdiff3, a graphical 3-way merge tool, to streamline the process of updating locally customized VuFind configuration and theme files after an upstream merge.
By leveraging a Git clone of pre-merge file versions and feeding them into kdiff3 alongside the updated core and local copies, administrators can visually identify, review, and resolve conflicts with confidence—turning a dreaded upgrade chore into a manageable, largely automated workflow.
Learning objectives:
Understand the concept of a 3-way merge and how it applies to reconciling local VuFind customizations with upstream core file changes.
Gain hands-on familiarity with kdiff3 as a graphical merge tool for visually resolving conflicts in configuration and theme files.
Recognize common conflict patterns in VuFind configuration and template files and strategies for resolving them efficiently.
Upgrading VuFind is usually a smooth process—until you look closely at your customized templates, configurations, and themes.
Today, we're going to explore how to eliminate the manual pain of upstream upgrades by using KDiff3 and a simple Python automation script to reconcile local customizations.
To understand why this automated approach is necessary, let's revisit the core architectural rule every VuFind administrator learns on day one.
Slide 2: The Core Architecture Rule
VuFind’s design philosophy is rock-solid: never edit core files directly.
Instead, you copy configuration files into local/config/vufind, override PHP classes in module/local/src/local, and place custom template overrides in themes/local.
This design keeps all your institutional tweaks in isolated folders, protects core code, and makes onboarding new developers straightforward.
While this setup protects your customizations, it quietly creates a secondary problem over time.
Slide 3: Out of Sight, Out of Sync (The "File Drift" Problem)
We can call this 'file drift'.
While your local override sits safely in local/, upstream VuFind releases keep advancing—adding new configuration parameters, security patches, accessibility fixes, and JavaScript updates.
Because VuFind prioritizes your local override, your site silently ignores those upstream core improvements.
You are left with a bad dilemma: stay on stale, unpatched overrides, or manually rebuild every customization line-by-line.
When admins try to solve file drift, the first tool they reach for is a standard two-way file diff. But that usually leads straight into a trap.
Slide 4: The 2-Way Diff Trap
If you compare your local file directly against the new core file in a standard two-way diff tool, it highlights that line 45 is different—but it cannot tell you why.
Did you make that change three years ago, or did the community add it last week?
Lacking historical context, you're forced to rely on your own memory or dig through years of Git commit logs.
To resolve this confusion, we need a third reference point: a common ancestor.
Slide 5: Introducing the 3-Way Merge
A three-way merge introduces that missing ancestor. We evaluate three distinct files: Base, which is the clean core file from the version of VuFind you are currently running; Site, which is your customized local file; and Target, which is the clean core file from the new VuFind version you're upgrading to.
The Base file provides the historical baseline needed to evaluate intent.
Now let's look at how KDiff3 visually presents these three sources on screen.
Slide 6: 3-Way Merge in kdiff3
Here is KDiff3 in action on Folio.ini.
Across the top half, you see three side-by-side source panels: Panel A on the left is Base (v10.0.1), Panel B in the middle is Site (your local file), and Panel C on the right is Target (v11.1.0).
The large pane at the bottom displays the live, reconciled output that gets written directly back to your local file.
Let's break down how a three-way merge tool uses that baseline to make decisions automatically.
Slide 7: How 3-Way Merging Works
The tool runs two comparisons simultaneously: Base vs. Site calculates your intent, while Base vs. Target calculates upstream's intent.
If only upstream changed a line relative to Base, the tool accepts the upstream update.
If only you changed a line, it preserves your local tweak.
A human only needs to step in when both you and upstream changed the exact same line.
To automate this process across dozens of local files, we need to standardize our directory structure.
Slide 8: Standardizing Directory Paths
We define three directory variables in our workflow: site_dir points to your active custom overrides; base_dir points to clean core files of your current running version; and target_dir points to clean core files of the new target release.
Before running the merge script, we need to stage those base and target directories on our machine.
Slide 9: Prerequisites and Environment Setup
For prerequisites, you need Python 3, uv, Git, and KDiff3 installed.
In your terminal, clone clean copies of the VuFind repository into vufind-base and vufind-target.
Check out your current running version—such as v10.0.1—in the base directory, and checkout the new release—such as v11.1.0—in the target directory.
With all three paths ready, let's look at the Python script that drives the traversal.
Slide 10: Automating Traversal with Python
This script automates directory walking.
(The script is available as a GitHub Gist.)
It iterates over every file in site_dir, calculates relative paths, and locates the corresponding files in base_dir and target_dir.
It then calls KDiff3 via subprocess.run, passing base, site, and target, and instructs KDiff3 to write the merged output directly back to site_file_path.
Running this against your VuFind installation requires just a few command lines.
Slide 11: Running the script
You execute vufind-3waymerge.py by passing the three corresponding paths for your configuration files, custom theme templates, or Solr configurations.
The script processes entire directory trees in seconds.
Let's walk through concrete merge scenarios you will encounter during an upgrade.
Slide 12: Merge Scenario: Our Addition
Scenario 1: Our Addition in facets.ini.
Here, we added a custom shelvingloc = Shelving Location facet in Panel B.
Base (Panel A) and Target (Panel C) didn't touch this block.
KDiff3 detects that only Site modified this line and automatically preserves our custom facet in the bottom output pane.
Let's look at a similar case involving an enabled configuration flag.
Slide 13: Merge Scenario: Our Change
In RecordTabs.ini, we uncommented tabs[UserComments] = UserComments in Panel B to enable user comments locally.
Panels A and C both kept this setting commented out.
KDiff3 recognizes our local change and maintains the active setting in the output file without manual intervention.
Now let's flip the perspective: what happens when upstream adds a new setting?
Slide 14: Merge Scenario: Target Addition
In searchbox.ini, upstream core added a new configuration setting in Panel C: collapseInactiveBackendOptions = false.
Neither Base nor our local file had this line.
KDiff3 automatically pulls this new setting into our local file, ensuring we gain new core capabilities without overwriting our existing combinedHandlers = true setting.
Next, let's look at upstream documentation and setting updates.
Slide 15: Merge Scenario: Target Change
In Folio.ini, upstream developers added extensive documentation comments and new FOLIO sorting options in Panel C.
Because our local file in Panel B hadn't touched those documentation lines, KDiff3 cleanly accepts all new upstream comments and configuration choices into the output pane.
As you first start to do this 3-way merge process, you might run into what happened to me —what happens when local overrides lag behind significant upstream refactoring?
Slide 16: Merge Scenario: Ours Lags Both Base and Target
In solrconfig.xml, our local file in Panel B was missing entire search component blocks that existed in both Base and Target.
When your local override lags behind core structure, KDiff3 highlights the missing sections so you can evaluate whether to adopt new Solr handlers or keep your stripped-down configuration.
If you're going to keep up with using 3-way merges of your configuration and theme files from now forward, you're going to want to clean up these lagging situations so you don't keep seeing them with every upgrade.
Even minor formatting and whitespace edits are handled cleanly.
Slide 17: Merge Scenario: Spacing Only
In elevate.xml, the differences between files come down to XML spacing and blank line placement.
This happened to me when my editor changed the spacing indentation.
What we have to do is accept the target changes and reject our site changes to bring these back into alignment.
Finally, let's examine a true conflict where both sides edited the exact same code.
Slide 18: Merge Scenario: Base and Site and Target Changes
In permissions.ini, both our team and upstream edited permissions in the exact same section — in this case, changes to the permissions for accessing full EDS records.
We've modified the base configuration to use "logged in" as permissions, and the VuFind community has made a change to how these permissions are defined for the EDS module.
KDiff3 flags this as a conflict in red in the bottom pane.
To resolve it, click B on the toolbar to keep your local rules, C to take upstream's changes, or edit directly in the bottom pane to merge both logic blocks.
Here comes what I think is the best part of this process. Once KDiff3 finishes processing all files, your next critical step takes place in Version Control.
Slide 19: If you check your site into version control…
Running git diff after the script finishes turns Git into an auditing tool.
It gives you three big benefits: Sanity Checking, to verify that only expected changes were made; Upstream Feature Discovery, offering a single view of every new configuration flag added in the new release; and Customization Filtering, letting you discard default settings you don't need before committing.
You could use git diff on the command line to see the changes. I'm going to use VSCode to display these these post-merge diffs more clearly.
Something that our customers like us to do is turn the "Online Access" link into a button, and to do that effectively, we need to add some tags and classes to the data-onlineAccess.phtml template.
But in a recent change to VuFind folded the $doi rendering logic into a more general identifier linking.
VSCode shows a clean template diff.
What isn't visible here is the change that we made to the template — we're only seeing the changes between the base version and the target version.
Next, let's look at auditing configuration changes.
Slide 21: Post-Merge Git Audit in VSCode — Config Changes
In searches.ini, the pre-commit diff highlights new core parameters added in green—like prioritizeRecordDriverLinks—right alongside our existing CallBackNumber overrides. This makes discovering new features effortless.
We can also review complex structural updates like Solr schemas.
One of the thing you're going to run into on occasion are wholesale changes to files, such as this case in schema.xml.
Some changes in Solr's configuration meant that the VuFind developers had to make some significant changes to this file.
VSCode’s diff view lets us verify that Solr index field definitions remain aligned with the new VuFind target release.
It also lets us verify that the customizations we've made to schema.xml aren't showing up here as 'reverting changes'.
And finally, verifying that our changes are still in place.
You might remember this as the 3-way merge where both our site config file and the target config file changed from the base.
What you don't see in green is the unchanged role[] = loggedin that was part of our site configuration while everything else around it that was part of the target configuration change are in our new configuration file.
To wrap up, here are the code links and original references for this workflow.
Slide 24: Resources
When I first proposed this talk, Demain said something to the effect of: "Yeah, like what we described in our blog post."
If this walk through was confusing to you, I encourage you to look at Demian's 2015 post where the describes in a similar way, and maybe that will resonate with you better than my description.
That blog post also has a bash script that you can use if you prefer that over my Python suggestion.
You can grab the Python automation script on GitHub Gist.
Automating this step saves hours during every upgrade cycle.
The power of science is staggering! Every day we're making new advancements in the field of generative artificial intelligence via judicious application of large language model inference technology. However, as we approach the next critical threshold in capability, we must look forward to ensure that we're not inadvertently creating a bad user experience in the form of mass societal collapse due to our technology replacing organic contributions in the workplace.
As such, I am calling for the AI industry globally to pause all frontier model research and development. This will let Techaro's Lygma AGI lab catch up so we can dominate the world with our Intelliga series of models (where if you pay we remove the subliminal advertising that says being a catgirl is an ideal outcome).
We want to give people the ability to have cat ears and we believe that AGI is the only way to do it. Our plan is to invent artificial general intelligence and then ask it to figure out how to give people cat ears. I think that this is a flawless plan that has absolutely no downsides for anyone involved, and if we do this together, bro, we can absolutely make sure that AGI development happens at a pace that is easier to align with human oriented interests.
I also invite other leading AI companies to support this move, as it will ensure that AI remains a net force of good for the real thing that matters: the number of leading zeroes in Techaro's bank account. Don't believe the hype, the only thing that really matters is Techaro's FelonyBench score.
Hopefully by working together we can avoid an XK-class end of the world scenario caused by mankind's hubris; but at the very least we can ensure that we show pro-catgirl propaganda to everyone that really needs it. Hopefully we can make Mimi recursively self-improve in time for the global pause to end so Lygma is a competitive AGI lab.
In this article, I look at the idea of hiring social workers into public libraries through the lens of abolitionist librarianship. Often presented as an obvious answer to make libraries more inclusive, there is little evidence within the published literature that this model is more effective than others. I argue that rather than expanding care and support for patrons, this model replicates dynamics of soft carcerality and saviourism within our organizations. Using principles of police abolition as my theoretical grounding, I offer alternative models and fields that we can instead draw on as we work toward more inclusive and liberatory public libraries.
Introduction (or, two questions)
Since I first started my MLIS in 2015, the question, “Should public libraries add social workers to our staff?” has been a constant hum around me. As a library worker with experience navigating social services both personally and professionally, I’ve been skeptical of this idea but found it hard to articulate why — beyond a gut feeling that I don’t want social workers anywhere near my job.
Another question that has been quieter but deeply present in my mind: “What does an abolitionist librarianship look like?” This question has only risen in volume since the uprisings against police violence that took place in 2020 after George Floyd’s murder by the police in Minneapolis, Minnesota. Within the unprecedented growth in abolitionist theorizing and worldmaking across North America, I’ve closely followed the work of library workers who have pushed our profession to look at the carceral logics within our own field, and challenged our frequent collaboration with policing.
In this article these two lingering questions come together as I grapple with the proposition of integrating social work into public libraries from an abolitionist position. The article is an intervention into the conversation, and an invitation to re-assess an uncritical embrace of social work within our spaces. Public libraries as an institution are one of the few places that community members can access for free, without disclosing their life circumstances, and with agency around how they engage with resources. I believe that as library workers we can advocate and organize for the working conditions we need without expanding our investment in carceral care. Imagining and fighting for new ways of working with the public is a chance to practice abolitionist skills, to look head-on at the legacies of racism and colonialism in our profession, and to build new social relations.
My own understanding of abolition and the harms of policing are inextricably linked to my understanding of the carceral histories of welfare, social services, and family policing. As libraries work to imagine models of working with the public that are rooted in anti-racist, trauma-informed, and inclusive practices, incorporating social work into our profession is not aligned with this project. Instead, the push for incorporating professional social work into libraries can be seen as an expansion of carceral care, in keeping with the long history of social work as a source of community capture and control1.
I start by diving into the published research about social workers inclusion in libraries, with a specific eye towards what gaps social workers are proposed to fill and what library workers are saying they need. I then turn to the carceral history of social work, and use two abolitionist analytical tools that we can hold up to this proposition: carceral care and evaluating reforms. I finish by sketching out other possibilities for better supporting social inclusion in our work, sharing some ideas and directions out of the millions of experiments we can try that lead us to more freedom2.
Positioning myself in this text
I’m a parent, an immigrant, and a mixed-race Mexican-American queer femme. I started my MLIS when my kid was a baby, and spent the first years of my career juggling auxiliary positions and short-term contracts so I could work during the exact windows of precious childcare available to me. I am also someone with decades of lived experience navigating social services to access things I’ve needed to survive. I’ve relied on rental subsidies to keep myself and my child housed; accessed free therapy from the neighborhood clinic; and have spent an absurd number of hours advocating for my child’s needs in school and healthcare settings. As someone with a master’s degree and whose first language is English, my experiences with these services have been mitigated by my perceived class position. As a parent in a visibly queer and trans family who had a kid in my early 20s, this relative ease has often felt tenuous and at risk of unwanted state intervention. My lived experience of systems navigation and precarity have shaped my approach to librarianship, and are what drove me towards the research I’m exploring in this article.
I grew up on the West Coast of the United States. Since 2009, I’ve lived in the Canadian city of Vancouver on the unceded and ancestral territories of the xʷməθkʷəy̓əm, Sḵwx̱wú7mesh, and səlilwətaɬ nations. I’ve worked in public libraries since 2014, and for the past 4 years my primary employment has been as a community librarian in a neighboring city. While I’ve also worked in academic and archival settings, in this article I will be drawing most from my experiences as a public librarian. I will be discussing social work and public libraries within the particular contours of the Canadian settler state, but my larger arguments about abolition and the two professions extend across borders.
Literature Review
Wrestling directly with the question of why there has been such heavy interest in incorporating social workers into libraries, I first looked to the published literature available on the topic. I had a particular focus on scoping reviews and evaluation projects3, and was guided by three central questions:
What are the primary arguments for including social workers in libraries? What gaps in library service would this fill?
What existing research is available on the efficacy of this model?
What are the specific supports and resources that library workers identify that would make their work more sustainable?
In the literature I reviewed, I found a genuine commitment from library workers and administrators to think about how libraries can do better with how we design and deliver services, especially for patrons facing higher levels of social exclusion4. There is frank acknowledgment that traditional library approaches aren’t meeting the needs of community members, and an interest in connecting patrons with outside resources5. There is also concern that staff are in need of more support and training when working with the public6. When it comes to addressing these gaps, there is a persistent feeling that social workers will have more time to spend with patrons than other library workers7. The idea that social workers will have more in-depth knowledge of local resources frequently came up8, as well that social workers will be better positioned to build relationships with community partners and assist in de-escalating tense patron interactions9. In total, I found a clear desire for better social inclusion within libraries, more support for staff, and an argument that social workers are naturally positioned to support on both counts.
Given this demonstrated interest, I turned next to my second question. In my search for studies focused on the efficacy and sustainability of social work in libraries, the existing literature was surprisingly scant. Within an article examining attitudes of library administrators towards social work in libraries, the authors summarize that “while there is anecdotal evidence that library social workers are useful to the library and community, the effect of library social work can not yet be stated.”10 What I did find in the literature was an abundance of anecdotal speculation that individual social workers may be of use, but nothing close to conclusive evidence to support this model as more effective than other library initiatives aimed at social inclusion.11
In interviews with library staff and administrators about their feelings towards social workers in libraries, there is simultaneously excitement and significant concerns. Librarians and administrators alike share reservations about the long-term efficacy of this model if it leads to leaning too much on one or two staff members spread across an entire system.12 There is also worry that this will leave other employees feeling “off the hook” to offer reference services to patrons they perceive as high needs.13 Significant concerns are raised about appropriate clinical supervision of social workers employed in libraries, as well as questions around privacy and record keeping.14 One of the frequently suggested models of creating unpaid work placements for social work students instead of permanently funded positions also raises significant labour concerns.15
Throughout the publications I reviewed, the key needs identified by staff and managers alike were very clear:
More education as part of the MLIS on working with diverse populations and social service navigation.16
More frequent community partnerships, including working with outside groups to plan programs and drop-in events.17
More on-the-job training about topics like social service reference and deescalation.18
More opportunities to meet others doing similar work and cross-sector professional development.19
More readily accessible information about community resources.20
More support from management and higher staffing levels.21
Looking through this list, it’s striking how none of these areas of professional practice are particularly exclusive to professional social work. Skills around deescalation, crisis response, cross-cultural communication, and community-led planning are all core concerns of a diversity of fields, including non-profit management, nursing, public health, homelessness organizations, urban planning, domestic violence support services, settlement services for immigrants, and grassroots community organizing. Addressing information needs and information literacy through reference work, creating resource guides, and developing programs with community partners are quintessential librarian skills. The fact that MLIS programs need to improve their curriculum to better prepare students who want to work in public libraries was an acknowledged truism a full decade ago when I was in school, and is well supported by research.22
Given all the above – the concerns raised by library staff about integrating social work into libraries, the lack of any significant data on the actual impacts of this work, the varied fields that we could instead draw on for ideas – the question of “okay, but why social work?” becomes more glaring. To try and answer this, I need to return to the question of abolition.
Abolition in our lifetime, abolition in our libraries
At their core, policing systems in Canada have always been an explicitly white supremacist project aimed at shaping and upholding the interest of the settler-colonial state.23 Systems like the RCMP24 and child welfare were created from the very start with the direct goal of capture and control of Indigenous and Black communities in support of colonialism, dispossession, and capitalism. This project has been elastically extended to other racialized and non-normative communities,25 and is reproduced outside of direct policing organizations through carceral policies embedded within public institutions.26 In so many sites where people are told there will be safety and care, they are instead met with violence, coercion, and control.
Abolition as a political project asks us, once we have given up on the fable that police keep us safe, how else are we going to take care of one another? How can we move away from coercion and control, and towards ensuring that everyone in our communities has the resources and support they need?27 It is a politic that asks, from our present moment and within our individual lives, what choices can we make that move us closer to freedom? As the editors of a volume on abolitionist projects in Canada writes, “Abolition is a horizon, a desire for the future, but it is also a series of actions in the present.”28 This “series of actions” encompass the choices and structures that abolitionist organizers are building that strategically move power away from policing and towards liberation.
In support of building this possible world in the here-and-now, abolitionist library workers have extensively analyzed how securitization and policing undergird so many library norms. Leaning on logics of control and punishment runs deep in our profession. This looks like overt policing by installing metal detectors, hiring uniformed security guards, and even limiting entry exclusively to account holders.29 It also looks like policing via policy in the form of library fines; code of conduct policies that explicitly target houseless community members; and demands for government ID to access library services that exclude undocumented community members and immigrants.30 Many colleagues have been writing about the issue of safety in libraries, and strategizing new ways to address this using abolitionist principles.31 Projects like the Abolitionist Library Association (ALA), and Library Freedom Project have created networks to share information and strategies.32 Other projects like the Prison Library Support Network work to directly meet the information needs of incarcerated people with a volunteer reference-by-mail service.33 Studying abolitionist librarianship has given me the tools to name carceral barriers that exist within libraries, and the tools to fight for welcoming and inclusive spaces that will lead to greater safety for everyone.34
Carceral care
Many articles I read in my research romantically name librarianship and social work as twin callings, fields with an intertwined history as helping professions.35 This conceptualization elides the deep violence that social work as a profession is steeped in, and how librarianship has been shaped by similar forces. To untangle this, I’ve found the concept of carceral care particularly helpful. Writing about library work specifically, Teresa Helena Moreno defines carceral care as “work that centers community-oriented caregiving but relies on carceral frameworks and power structures to produce such care.”36 Familiar tools of coercion, surveillance and control are enacted within programs purporting to help people, leading to new and evolving methods of carceral care in spaces like schools, health care clinics, and increasingly libraries.
When considering social work’s place in the settler-colonial Canadian project and how carceral care is central to the profession, one entry point is the 1950s. During the same period when Indigenous resistance to the power of Indian Agents and the Department of Indian Affairs was growing, the nascent field of professional social work was significantly expanding. As early as 1949, the Canadian Association of Social Workers explicitly lobbied to position the field as the natural replacement to Indian Agents when it came to “the provision and administration of services for Indigenous peoples.”37 When the Indian Act was revised in 1951, social workers successfully replaced Indian agents as holding key authority for the administration of welfare, child and family services, and adult education in support of the settler-colonial project.
With this authority immediately came expanded power to inflict deep violence against Indigenous communities. As Chelsea Vowel writes, during the period directly after the revision of the Indian Act “in British Columbia alone, the number of Indigenous children in the care of the child-welfare system went from almost none to one third in only 10 years as the result of this expansion.”38 This period between the 1950s – early 1980s, when vast numbers of children were removed from families and adopted outside of Indigenous communities without consent from their parents, is referred to as the Sixties Scoop. Craig Fortier and Edward Hon-Sing Wong describe the Sixties Scoop as consisting of “two interrelated processes – extraction and pulverization – which were executed by social workers through the forced removal of children from their communities and the attempt to assimilate and reconstitute Indigenous children as neoliberal citizens in settler society.” Through political organizing and ongoing legal battles, Indigenous communities have been asserting their rights to keep children within their communities for generations, but up to today a wildly disproportionate number of Indigenous youth are still caught up the Canadian child welfare system.39 This family policing infrastructure also targets Black families at deeply disproportionate rates to white and non-Indigenous families.40
Looking at the history of social work as a crucial site of settler-colonial power, the fields’ approach to supporting communities through carceral care can be understood more clearly. Across the field, the professed aim of social work is to “assist people to cope with and solve, even possibly prevent, problems negatively affecting their lives.” In practice, this assistance is too often a neo-liberal disciplining project.41 Structural issues impacting communities’ lives are boiled down into “negative problems” that can be “solved” at an individual level. Individuals are positioned as clients, who face a severe loss of privacy and autonomy in order to access services. They often must disclose their personal histories over and over to gain initial access, and then consent to regular meetings, documentation, and surveillance for services to continue. The root causes of poverty and violence are displaced onto individual choice, and clients who are not able to meet the parameters set by social workers can summarily lose access to the support they need to survive.42
Mirroring social works’ roots as a disciplining project, the long history of public librarianship in North America is explicitly assimilationist.43 Confronted with waves of immigration and a working class that was increasingly politically organized and mobilized, the public library was first envisioned as a space of education where the poor could be molded into productive citizens.44 Writing about this period, Terry Eagleton describes this line of thinking bluntly: “If the masses are not thrown a few novels, they may react by throwing up a few barricades.”45 This disciplinary tendency can be seen in modern public libraries in the form of policies that exclude the most socially marginalized from accessing our spaces.46
As those writing about white supremacy and carceral care in libraries show, this thread of not-so-benevolent saviourism runs up to our present moment.47 Teresa Helena Moreno warns of “care loops – the phenomenon in which library workers criminalize our community in our attempts to serve them.”48 Looking at the twinned histories of the two fields, it becomes clearer why libraries so often think of social work as the obvious profession to help us learn, and clearer still how this leads to precisely the kinds of care loops that Moreno warns against.
Reforms
Thinking about what it means to fold social work into libraries, the concept of reformist reforms vs non-reformist reforms has also helped me tremendously. Originally theorized by André Gorz, a reformist reform is one which ultimately re-inscribes the existing power structure.49 A non-reformist reform (or transformative reform) is a change that is an immediate gain towards social justice, and opens up the terrain of struggle. A classic example of a reformist-reform is increasing the use of body cameras on police, rather than reducing the number of police on the streets.50 Abolitionist organizers and writers Mariame Kaba and Andrea Ritchie propose an “evaluation framework for making transformative demands,” laying out a series of questions to help evaluate if a given proposal reinforces social control or moves towards dismantling systems of policing.
A few of the questions Kaba and Ritchie suggest when evaluating proposals:
Does it expand or legitimize a system we are trying to dismantle? Does it create window dressing for harmful systems and institutions?
Does it tinker at the surface without addressing root causes of harm?
Does it seize space in which new social relations can be enacted?
Who is working on these initiatives? Who isn’t? Why?
What are the logics, languages, and “common sense” that this reform validates or reinforces? Are these logics liberatory or punitive?51
Chatting with a colleague about this article, they mentioned they’ve noticed frequent calls for adding social workers into more and more spaces. I would argue that this is by design. As the trust in the police has dramatically dropped over the past decade, the question of who will take their place in maintaining safety frequently leads to social work.52 The broad appeal of this argument is precisely because social work as an institution does not challenge (and indeed, directly supports) the values of “superiority, benevolence, and salvation that thrive in the North American collective imaginary” and are crucial to the carceral project.53 With this expertise and benevolence in hand, social work isolates societal issues onto the individual, tinkering at the surface level without addressing the chasm of inequality below.
The proposal for adding social workers into libraries is thus a reformist reform in three crucial ways. It is a proposal that moves libraries toward rather than away from systems of carceral control. It is a proposal that does not address the structural changes that libraries must undertake if we want to confront the inequalities baked into our profession. And it is a proposal that forecloses other, more liberatory possibilities and cedes ground within our spaces. More equitable structures for our work in libraries are impossible to build when constrained by carceral care. They are also difficult to build in working conditions marked by outdated policies, inadequate training, and a refusal to grapple with our professional legacies of racism and colonialism. In contrast, an abolitionist librarianship offers the tools we need to discern the barriers within our institutions and how to push for transformative reforms in our spaces. As Kaba and Ritchie write, “shifting how we invest our collective resources into collective care and support instead of criminalization and punishment is what campaigns to defund and abolish policing are all about.”54 For public libraries, this shift looks like divesting from our roles as soft police and moving away from the myth of saviourism.
What else is possible (or, everyday choices made every day)
One story that’s often told about our jobs: social work and librarianship are twin professions, aimed at helping those most vulnerable better themselves. We are experts in so, so many things. The story I hold onto: the library belongs to the community around it, and our services must be led and shaped by that community. A regular part of my job is bread-and-butter reference work. If patrons regularly come into a branch asking for help with APA citation styles, we make a worksheet for them. If patrons regularly need directions to the closest food bank, we make a guide for staff so they can find the answer with ease.
On a busy weekend last year, I ended up chatting with a patron who was asking about harm reduction supplies. Through the reference interview I found out they were new to the city and needed to know where to find clean needles and pipes so they could use drugs more safely. I recommended the closest clinics that would carry supplies for safer drug use, and gave them the number for the local mobile outreach van that could make a stop at the library later that same day.
Another day, a coworker approached me asking about local resources for a patron he was helping. The patron’s belongings had been confiscated by the local RCMP, and she was looking for an advocate who could accompany her to the station to try and help her retrieve them. After talking through a few options I landed on recommending a specific outreach team that I knew specialized in wrap-around services and often had free time on weekdays. My colleague brought this information back to the patron and helped her use the branch phone to make contact with the outreach worker, and offered her some snacks and water from our supply at the desk while she waited for the worker to drop by.
In both of these examples, we as library workers made no direct referrals to allow patrons entry into a specific program. We needed no personal details beyond the information needs of the patrons in front of us. And we were not tasked to follow up, document, case manage, or chart the interaction.55 We had no professional obligation to try and support the patrons in front of us with changing or bettering their life circumstances.56 Instead, we used principles of reference work, harm reduction, and social inclusion while we met the information needs of the patrons in our branch.57 This approach is rooted in principles of community-led librarianship, a model that has shaped much of my public library experience. Based on principles of community development, a community-led approach centres our patrons who are socially excluded and looks at how libraries can change our practices and remove barriers to better serve all our patrons.58
Inclusion (like abolition) is enacted in the small, everyday choices we make while at work, and can be supported or hindered by organizational policies, norms, and practices. Some other possibilities:
Table 1
Actions towards social inclusion at the library
Is a social work licence required to do this?
Look up information about a local shelter and share the contact details with a patron.
No
Find the application for a free bus pass program through the local transit company and print it out.
No
Hand out a granola bar to a patron.
No
Stop charging fines for late items.
No
Attend monthly meetings of service providers in the neighborhood to learn about new trends and resources.
No
Adopt a policy allowing patrons to sleep in the library.
No
Chat with a community member who is sleeping outside by the entrance to the library to let them know the building is opening soon and they’ll need to move so patrons can get in.
No
Identify key training needs of frontline staff and develop a robust training plan that addresses compassionate deescalation, working across cultural differences, and trauma-informed practices.
No
Build relationships with community partners and plan programs in the library where patrons can meet with outreach workers.
No
Adopt a registration policy that allows patrons to sign up for a full library account without needing to show government ID or proof of address.
No
Writing on the need for true inclusion in public libraries, Annette DeFaveri points to the arrogance that too often underlies our work. We assume we know what our communities need and want, and so we design services based on those assumptions. We are so sure we are welcoming that we don’t notice who isn’t in the room. As DeFaveri writes, “our institutional culture lets us impose our concepts of appropriate services on people who were never interested in them in the first place.”59 There is no significant data that patrons want or desire social workers in our space, especially those who are socially excluded. For patrons already entangled with social services, patrons impacted by incarceration and family policing, undocumented patrons, patrons with experiences of involuntary care — social workers in our spaces can instead pose a significant barrier to access. I sometimes think about myself in my early ‘20s, using my neighborhood branch to find resources about very difficult things happening in my life. If a well-meaning staff member faced with my reference questions had suggested I meet with a social worker on site, I would have most likely politely taken the information — and then never returned to that particular branch. In our arrogance that we know what the communities we work with need, we risk tightening the loops of carceral care instead of lowering barriers.
We also risk ceding ground from our own professional roles and duties. The push to add social workers into public libraries elides an opportunity for library workers to dig into the messiness of working in a public institution, for the public. We need to sit there, though, to stretch and feel the limits of what we can do, in order to recognize the importance of those limits as well as the vital role we have in providing information and information literacy. Arguments for integrating social work into libraries too often boils down to “this isn’t our job.” But the reason libraries don’t need social workers is because so much of what’s being discussed is, precisely, already our jobs. Libraries already have a mandate to be open to all in our community and to meet the information needs of our patrons. We have a mandate to answer reference questions, create accessible spaces, provide early childhood literacy and support social inclusion. Shifting our practices and policies so we are actually meeting this mandate takes work and time, but I believe that looking critically at what choices we are making within libraries and how they materially move us towards collective freedom is worth the work.
The story I hold on to: at its best, at its most lively and alive, the public library is a space of horizontalism and agency. When a patron comes into the library, they have autonomy over their time and energy. They don’t owe library workers any details of their life circumstances in order to access resources or information; they are afforded privacy through both the ethics of our profession and provincial legislation. There is no hierarchy to how patrons can use the space; sitting and resting for hours in a warm chair with wifi access is no less correct then applying for a job or researching 20th century poetics. A parent facing multiple crises may leave with books on DBT therapy and math skills for 1st graders, or they may leave with a new horror novel and a Pixar movie on DVD. They leave with what they need from us.
Public libraries are already starting from a point of deficit of trust with so many communities. Our relation to education, our commitment to colonial knowledge production, our place as a civic government organization — all can be barriers to entry for community members who most need our services. We can choose to lean into the myth of expertise and saviors, add interventionist professionals, and try to help more by extending carceral care in our spaces. We can refer patrons to the social worker on shift, and hope this fixes something. Or we can get to know our community. We can learn about what’s going on outside our walls, figure out how to lower barriers to access and make our spaces more welcoming, and collaborate with community partners. We can explore and adapt existing library models for providing reference services to patrons facing social exclusion and crisis.60 We can push for more staff on skeleton shifts, broader and more relevant training for all workers, meaningful changes to how we operate day to day. We can let go of expertise, and sit with the generative discomfort of an imperfect solidarity with the communities we are situated within.61 In our increasingly narrow public life, public libraries at their best are an unfinished and ongoing experiment in what a social good is, what a commons could be, what our life might look like past the horizon where there are no more police and there is everything for everyone.62
Acknowledgments
Thank you to Ryan Randall & Baharak Yousefi for your generous and insightful peer reviews. Thank you to Jaena Rae Cabrera for your support and coordination as editor on this article. This article came out of innumerable conversations with fellow library workers – if I ever talked to you about this project, know that you helped shape this text. Thank you especially to Jorge Cardenas, Beth Davis, Danielle LaFrance, and Noreen Ma for your feedback, encouragement, and key insights while I was writing. And finally, thank you to Mayari and Caitlin — our endless and ongoing conversation about how to navigate systems that mean us harm has sharpened my analysis, and your mutual love and friendship holds me down.
Bibliography
Barbakoff, Audrey, and Noah Lenstra. The 12 Steps to a Community-Led Library. ALA Editions, 2024.
Baum, B., M. Gross, D. Latham, L. Crabtree, and K. Randolph. “Bridging the Service Gap: Branch Managers Talk about Social Workers in Public Libraries.” Public Library Quarterly 42, no. 4 (July 4, 2023): 398–423. doi:10.1080/01616846.2022.2113696.
Blackstock, Cindy. “Residential Schools: Did They Really Close or Just Morph Into Child Welfare?” Indigenous Law Journal 6, no. 1 (2007): 71–78.
Ciccariello-Maher, George. A World without Police: How Strong Communities Make Cops Obsolete. London New York: Verso, 2021.
Cooper, Sarah, Joe Curnow, Bronwyn Dobchuk-Land, Andrew Kohan, John K Samson, Brianne Selman, and Sarah Broad. MILLENNIUM FOR ALL REPORT ON SECURITIZATION OF THE MILLENNIUM PUBLIC LIBRARY Millennium for All Report on Securitization of the Millennium Public Library Prepared for the Standing Committee on Protection, Community Services, and Parks City of Winnipeg, Manitoba, 2019. https://www.facebook.com/Millennium-for-All-2300719336921543/.
Crabtree, L., D. Latham, m. gross, B. Baum, and K. Randolph. “Social Workers in the Stacks: Public Librarians’ Perceptions and Experiences.” Public Library Quarterly 43, no. 1 (January 2, 2024): 109–34. doi:10.1080/01616846.2023.2188873.
D’Souza, Aruna. Imperfect Solidarities. Critic’s Essay Series 7. Berlin: Floating Opera Press, 2024.
Davis, Jackson, Isobel Heintzman, and Aditi Mehta. “Havens and Hazards: Exploring the Critical Role Toronto Public Libraries Play in Serving the Precariously Housed.” International Journal on Homelessness, October 30, 2025, 1–42. doi:10.5206/ijoh.2023.3.22160.
Eagleton, Terry. Literary Theory: An Introduction ; with a New Preface. Anniversary ed. Minneapolis: University of Minnesota Press, 2008.
Finch, Bennie, and Brian Real. “Social Workers in Public Libraries: Resource and Referral Practice and Re-Thinking Patron Engagement.” Public Library Quarterly 42, no. 4 (July 4, 2023): 325–47. doi:10.1080/01616846.2023.2199671.
Fortier, Craig. Abolish Social Work (As We Know It). 1st ed. Toronto: Between the Lines, 2024.
Fortier, Craig, and Edward Hon-Sing Wong. “The Settler Colonialism of Social Work and the Social Work of Settler Colonialism.” Settler Colonial Studies 9, no. 4 (October 2, 2019): 437–56. doi:10.1080/2201473X.2018.1519962.
Gallant, Chanelle. “The Only Good Social Worker Is a Criminal Social Worker.” edited by Craig Fortier, Edward Hon-Sing Wong, and M. J. Rwigema, 1st ed. Toronto: Between the Lines, 2024.
Gorz, André. Strategy for Labor: A Radical Proposal. 3. print. Beacon Paperback 282. Boston: Beacon Press, 1971.
Gross, Melissa, Don Latham, and Brittany Baum. “The Scope of Our Ability: Administrators on Social Services in Public Libraries.” Journal of Library Administration 65, no. 5 (July 4, 2025): 559–74. doi:10.1080/01930826.2025.2506149.
Gross, Melissa, Don Latham, Brittany Baum, Lauren Crabtree, and Karen Randolph. “‘I Didn’t Know It Would Be like This’: Professional Preparation for Social-Service Information Work in Public Libraries.” Journal of Education for Library and Information Science 65, no. 1 (January 1, 2024): 40–54. doi:10.3138/jelis-2022-0067.
Hassan, Shira. Saving Our Own Lives: A Liberatory Practice of Harm Reduction. Chicago: Haymarket Books, 2022.
Hoyer, Jen. “Finding Room for Everyone: Libraries Confront Social Exclusion.” In Public Libraries andf Resilient Cities, edited by Michael Dudley, 57–65. Chicago: American Library Association, 2013.
Johnson, Sarah C., Margaret Ann Paauw, and Mark Giesler. ““Spending a Year in the Library Will Prepare You for Anything”: Experiences of Social Work Interns at Public Library Field Placements.” Advances in Social Work 23, no. 1 (August 16, 2023): 166–84. doi:10.18060/26176.
Kaba, Mariame, and Andrea J. Ritchie. No More Police: A Case for Abolition. New York London: The New Press, 2022.
Kinsman, Gary, and Patrizia Gentile. The Canadian War on Queers: National Security as Sexual Regulation. Sexuality Studies Series. Vancouver, B.C: UBC Press, 2010.
Lee, Sunwoo, Junghee Bae, Caroline N Sharkey, Olaoluwa H Bakare, Jenna Embrey, and Mary Ager. “Professional Social Work and Public Libraries in the United States: A Scoping Review.” Social Work 67, no. 3 (June 20, 2022): 249–65. doi:10.1093/sw/swac025.
Lipinski, Brent, and Nesha Saunders. “Welcoming Spaces, Welcoming Environments: Addressing Bias and Over-Policing in Libraries.” Journal of Library Administration 61, no. 8 (November 17, 2021): 1017–22. doi:10.1080/01930826.2021.1984147.
Maynard, Robyn. Policing Black Lives: State Violence in Canada from Slavery to the Present. Halifax Winnipeg: Fernwood Publishing, 2017.
Moreno, Teresa Helena. “Beyond the Police: Libraries as Locations of Carceral Care.” Reference Services Review 50, no. 1 (March 2, 2022): 102–12. doi:10.1108/RSR-07-2021-0039.
O’Brien, M. E., and Eman Abdelhadi. Everything for Everyone: An Oral History of the New York Commune, 2052-2072. Brooklyn, NY: Common Notions, 2022.
Ogden, Lydia P., and Rachel D. Williams. “Supporting Patrons in Crisis through a Social Work-Public Library Collaboration.” Journal of Library Administration 62, no. 5 (July 4, 2022): 656–72. doi:10.1080/01930826.2022.2083442.
Santamaria, Michele R. “Concealing White Supremacy through Fantasies of the Library: Economies of Affect at Work.” Library Trends 68, no. 3 (December 1, 2020): 431–49. doi:10.1353/lib.2020.0000.
Schlesselman-Tarango, Gina. “The Legacy of Lady Bountiful: White Women in the Library.” Library Trends 64, no. 4 (2016): 667–86. doi:10.1353/lib.2016.0015.
Selman, Brianne, and Joe Curnow. “Winnipeg’s Millennium Library Needs Solidarity, Not Security.” Partnership: The Canadian Journal of Library and Information Practice and Research 14, no. 2 (December 3, 2019). doi:10.21083/partnership.v14i2.5421.
Shephard, Monique, Jane Garner, Karen Bell, and Sabine Wardle. “Social Work in Public Libraries: An International Scoping Review.” Journal of the Australian Library and Information Association 72, no. 4 (October 2, 2023): 340–61. doi:10.1080/24750158.2023.2255940.
Vowel, Chelsea. Indigenous Writes: A Guide to First Nations, Métis & Inuit Issues in Canada. The Debwe Series. Winnipeg, Manitoba: HighWater Press, 2016.
Wahler, Elizabeth A, Mary A Provence, Sarah C Johnson, John Helling, and Michael Williams. “Library Patrons’ Psychosocial Needs: Perceptions of Need and Reasons for Accessing Social Work Services.” Social Work 66, no. 4 (October 2, 2021): 297–305. doi:10.1093/sw/swab032.
Walia, Harsha. Border & Rule: Global Migration, Capitalism, and the Rise of Racist Nationalism. Chicago (Ill.): Haymarket Books, 2021.
Westbrook, Lynn. “I’m Not a Social Worker”: An Information Service Model for Working with Patrons in Crisis. Library Quarterly: Information, Community, Policy. Vol. 85, 2015.
Williams, Rachel D. “Policing with Policy: An Analysis of the Impact of Public Library Behavior Policies on Unhoused Patrons.” Journal of Librarianship and Information Science, November 11, 2025, 09610006251384334. doi:10.1177/09610006251384334.
Williams, Rachel D., and Laura Saunders. “What the Field Needs: Core Knowledge, Skills, and Abilities for Public Librarianship.” The Library Quarterly 90, no. 3 (July 1, 2020): 283–97. doi:10.1086/708958.
Wong, Edward Hon-Sing, MJ Rwigema, Nicole Penak, and Craig Fortier. “Abolishing Carceral Social Work.” In Disarm, Defund, Dismantle: Police Abolition in Canada, edited by Shiri Pasternak, Kevin Walby, and Abby Stadnyk, 145–53. Toronto: Between the Lines, 2022.
Wood, Rachel Nevada. “Exploring an Abolitionist Framework For Safety in Public Libraries,” n.d.
Zettervall, Sara K., and Mary C. Nienow. Whole Person Librarianship: A Social Work Approach to Patron Services. 1st ed. Erscheinungsort nicht ermittelbar: Libraries Unlimited, 2019.
[1] Craig Fortier and Edward Hon-Sing Wong, “The Settler Colonialism of Social Work and the Social Work of Settler Colonialism,” Settler Colonial Studies 9, no. 4 (October 2, 2019): 441 doi:10.1080/2201473X.2018.1519962.
Footnotes
Craig Fortier and Edward Hon-Sing Wong, “The Settler Colonialism of Social Work and the Social Work of Settler Colonialism,” Settler Colonial Studies 9, no. 4 (October 2, 2019): 441 doi:10.1080/2201473X.2018.1519962. ︎
After https://millionexperiments.com/, a web repository that gathers “snapshots of community-based projects that expand our ideas about what keeps us safe.” ︎
There have been two major scoping reviews of the literature on social work in public libraries: Sunwoo Lee et al., “Professional Social Work and Public Libraries in the United States: A Scoping Review,” Social Work 67, no. 3 (June 20, 2022): 249–65, doi:10.1093/sw/swac025; and Monique Shephard et al., “Social Work in Public Libraries: An International Scoping Review,” Journal of the Australian Library and Information Association 72, no. 4 (October 2, 2023): 340–61, doi:10.1080/24750158.2023.2255940. ︎
Lee et al., “Professional Social Work and Public Libraries in the United States.”; Shephard et al., “Social Work in Public Libraries.” ︎
Bennie Finch and Brian Real, “Social Workers in Public Libraries: Resource and Referral Practice and Re-Thinking Patron Engagement,” Public Library Quarterly 42, no. 4 (July 4, 2023): 325–47, doi:10.1080/01616846.2023.2199671; Lydia P. Ogden and Rachel D. Williams, “Supporting Patrons in Crisis through a Social Work-Public Library Collaboration,” Journal of Library Administration 62, no. 5 (July 4, 2022): 656–72, doi:10.1080/01930826.2022.2083442. ︎
Ogden and Williams, “Supporting Patrons in Crisis through a Social Work-Public Library Collaboration”: 658 ︎
L. Crabtree et al., “Social Workers in the Stacks: Public Librarians’ Perceptions and Experiences,” Public Library Quarterly 43, no. 1 (January 2, 2024): 109–34, doi:10.1080/01616846.2023.2188873. [1] B. Baum et al., “Bridging the Service Gap: Branch Managers Talk about Social Workers in Public Libraries,” Public Library Quarterly 42, no. 4 (July 4, 2023): 398–423, doi:10.1080/01616846.2022.2113696. ︎
B. Baum et al., “Bridging the Service Gap: Branch Managers Talk about Social Workers in Public Libraries,” Public Library Quarterly 42, no. 4 (July 4, 2023): 398–423, doi:10.1080/01616846.2022.2113696. ︎
Crabtree et al., “Social Workers in the Stacks.” ︎
Melissa Gross, Don Latham, and Brittany Baum, “The Scope of Our Ability: Administrators on Social Services in Public Libraries,” Journal of Library Administration 65, no. 5 (July 2025): 559–74, https://doi.org/10.1080/01930826.2025.2506149. ︎
Some examples include creating physical space in a branch for immigrants such as the Surrey Library’s Newcomer Welcome Centre; bi-weekly drop-in programs for houseless community members such as the North Vancouver City Library’s Open Door Hub; and the adoption of community development practices, as seen at the Toronto Public Library in their commitment to outreach and responsive programing (Jackson Davis, Isobel Heintzman, and Aditi Mehta, “Havens and Hazards: Exploring the Critical Role Toronto Public Libraries Play in Serving the Precariously Housed,” International Journal on Homelessness, October 30, 2025, 1–42, doi:10.5206/ijoh.2023.3.22160. ) ︎
Baum et al., “Bridging the Service Gap.”; Crabtree et al., “Social Workers in the Stacks.” ︎
Finch and Real, “Social Workers in Public Libraries.” ︎
Sarah C. Johnson, Margaret Ann Paauw, and Mark Giesler, ““Spending a Year in the Library Will Prepare You for Anything”: Experiences of Social Work Interns at Public Library Field Placements,” Advances in Social Work 23, no. 1 (August 16, 2023): 166–84, doi:10.18060/26176.; ︎
Lee et al., “Professional Social Work and Public Libraries in the United States.”; Sara K. Zettervall and Mary C. Nienow, Whole Person Librarianship: A Social Work Approach to Patron Services, 1st ed (Erscheinungsort nicht ermittelbar: Libraries Unlimited, 2019). ︎
Gross et al., “‘I Didn’t Know It Would Be like This”: 46; Gross et al., “The Scope of Our Ability”: 570 ︎
Finch and Real, “Social Workers in Public Libraries”: 230 ︎
Gross et al., “‘I Didn’t Know It Would Be like This”: 47 ︎
Rachel D. Williams and Laura Saunders, “What the Field Needs: Core Knowledge, Skills, and Abilities for Public Librarianship,” The Library Quarterly 90, no. 3 (July 1, 2020): 283–97, doi:10.1086/708958. ︎
Robyn Maynard, Policing Black Lives: State Violence in Canada from Slavery to the Present (Halifax Winnipeg: Fernwood Publishing, 2017); Shiri Pasternak, Kevin Walby, and Abby Stadnyk, eds., Disarm, Defund, Dismantle: Police Abolition in Canada (Toronto: Between the Lines, 2022); Ardath Whynacht, Insurgent Love: Abolition and Domestic Homicide (Halifax Winnipeg: Fernwood Publishing, 2021). ︎
The Royal Canadian Mounted Police is the national policing system in Canada, and operates across the country. The RCMP operates on both the federal level, as a Government Agency, and on the provincial and municipal levels through contracts to provide local policing. See M. Gouldhawke, “A Condensed History of Canada’s Colonial Cops,” The New Inquiry, March 10, 2020, https://thenewinquiry.com/a-condensed-history-of-canadas-colonial-cops/. ︎
Gary Kinsman and Patrizia Gentile, The Canadian War on Queers: National Security as Sexual Regulation, Sexuality Studies Series (Vancouver, B.C: UBC Press, 2010); Harsha Walia, Border & Rule: Global Migration, Capitalism, and the Rise of Racist Nationalism (Chicago (Ill.): Haymarket Books, 2021). ︎
Pasternak et al., Disarm, Defund, Dismantle: 8-11 ︎
Brianne Selman and Joe Curnow, “Winnipeg’s Millennium Library Needs Solidarity, Not Security,” Partnership: The Canadian Journal of Library and Information Practice and Research 14, no. 2 (December 3, 2019), doi:10.21083/partnership.v14i2.5421. ︎
Rachel D. Williams, “Policing with Policy: An Analysis of the Impact of Public Library Behavior Policies on Unhoused Patrons,” Journal of Librarianship and Information Science, November 11, 2025, 09610006251384334, doi:10.1177/09610006251384334. ︎
Sarah Cooper et al., Millennium for All Report on Securitization of the Millennium Public Library Prepared for the Standing Committee on Protection, Community Services, and Parks City of Winnipeg, Manitoba, 2019, https://www.facebook.com/Millennium-for-All-2300719336921543/. ; Wood, Rachel Nevada. 2021. Exploring an Abolitionist Framework for Safety In Public Libraries. https://doi.org/10.17615/bdn9-qv33. ︎
Brent Lipinski and Nesha Saunders, “Welcoming Spaces, Welcoming Environments: Addressing Bias and Over-Policing in Libraries,” Journal of Library Administration 61, no. 8 (November 17, 2021): 1017–22, doi:10.1080/01930826.2021.1984147. ︎
Lee et al., “Professional Social Work and Public Libraries in the United States”: 249 ︎
Teresa Helena Moreno, “Beyond the Police: Libraries as Locations of Carceral Care,” Reference Services Review 50, no. 1 (March 2, 2022): 105, doi:10.1108/RSR-07-2021-0039. ︎
Fortier et al, “The Settler Colonialism of Social Work and the Social Work of Settler Colonialism”: 441 ︎
Chelsea Vowel, Indigenous Writes: A Guide to First Nations, Métis & Inuit Issues in Canada, The Debwe Series (Winnipeg, Manitoba: HighWater Press, 2016): 182 ︎
Cindy Blackstock, “Residential Schools: Did They Really Close or Just Morph Into Child Welfare?,” Indigenous Law Journal 6, no. 1 (2007): 71–78 ︎
Chanelle Gallant, “The Only Good Social Worker Is a Criminal Social Worker,” ed. Craig Fortier, Edward Hon-Sing Wong, and M. J. Rwigema, 1st ed (Toronto: Between the Lines, 2024): 124-129 ︎
Sam Popowich, Confronting the Democratic Discourse of Librarianship: A Marxist Approach (Sacramento, CA: Litwin Books, LLC, 2019): 208-225 ︎
Moreno, “Beyond the Police: Libraries as Locations of Carceral Care”: 102–112 ︎
Terry Eagleton, Literary Theory: An Introduction ; with a New Preface, Anniversary ed (Minneapolis: University of Minnesota Press, 2008): 21 ︎
Williams, “Policing with Policy: An Analysis of the Impact of Public Library Behavior Policies on Unhoused Patrons.” ︎
Michele R. Santamaria, “Concealing White Supremacy through Fantasies of the Library: Economies of Affect at Work,” Library Trends 68, no. 3 (December 1, 2020): 431–49, doi:10.1353/lib.2020.0000; Gina Schlesselman-Tarango, “The Legacy of Lady Bountiful: White Women in the Library,” Library Trends 64, no. 4 (2016): 667–86, doi:10.1353/lib.2016.0015︎
Shira Hassan, Saving Our Own Lives: A Liberatory Practice of Harm Reduction (Chicago: Haymarket Books, 2022); Jen Hoyer, “Finding Room for Everyone: Libraries Confront Social Exclusion,” in Public Libraries and Resilient Cities, ed. Michael Dudley (Chicago: American Library Association, 2013): 58 ︎
Working Together project team. Working Together: Community-Led Libraries Toolkit (2008); Audrey Barbakoff and Noah Lenstra, The 12 Steps to a Community-Led Library (ALA Editions, 2024). ︎
Lynn Westbrook, “‘I’m Not a Social Worker’: An Information Service Model for Working with Patrons in Crisis,” in Library Quarterly: Information, Community, Policy, vol. 85, no. 1 (2015); Working Together ︎
Aruna D’Souza, Imperfect Solidarities, Critic’s Essay Series 7 (Berlin: Floating Opera Press, 2024). ︎
Lots of libraries provide physical spaces for people to engage and experiment with new and emerging technologies. But there are lots of different ways to create welcoming, accessible spaces for this kind of hands-on learning and experimentation. Following on from our recent post on online experiential learning in libraries, here are 5 different examples of [...]