Requirements from a DH perspective

From DHWiki

This is the description of the methodology applied for collecting software engineering requirements for the DHwiki white paper. The expected readers are digital humanities practitioners to validate each entry and software engineers to implement the features. For the non technical savvy practitioners, we try to make the descriptions as less obfuscated as possible to be easily validated. but with enough detail to be useful for the engineers.

Disclosure: parts of the text here has been created with the assistance of ChatGPT,[1] under the direction, scrutiny and review of the editor.

Definition

The SWBOK[2] defines a requirement as:

a condition or capability needed by a user to solve a problem or achieve an objective;“
Second paragraph.“

And the ISO/IEC/IEEE 24765 vocabulary[3] defines it as:

condition or capability that must be met or possessed by a system, system component, product, or service to satisfy an agreement, standard, specification

Sources

The information has been collected by an editor but the primary sources are many:

  • Selected issues tracked in https://phabricator.wikimedia.org/ related with Wikibase.
  • Direct testimonials of persons with direct involvement in a Wikibase implementation in a digital humanities project.
  • Reports and documentation from digital humanities Wikibase projects.
  • And the own experience implementing the knowledge graph for the Centro Documental del Cine de Almería.

Use case

Use case for the compilation of software requirements for the digital humanities.

For scoping the compilation of requirements we designed a use case. A UML use case[4] is a simple way to describe what a system is supposed to do from the point of view of the people who use it. Instead of talking about code or technical details, it focuses on the objectives: who interacts with the system (the “actors”, such as users or other systems) and what they want to achieve (the “use cases”, such as “submit a request” or “search the archive”). A use case diagram is like a map of responsibilities and interactions: it shows the main services offered by the system and which groups depend on them, helping everyone agree on the scope and expectations before any construction begins.

Description

This diagram tells a simple story about how to turn needs into clear, traceable software requirements, and then how engineers use those requirements to build software.

Who is involved

Four groups of people take part:

  • DHwiki Members: people involved in the DHwiki initiative/community.
  • Digital Humanities practitioners: researchers and practitioners who will use the tools and can explain what they need.
  • Wikimedia Developers: the wider Wikimedia technical community who can contribute ideas, constraints, and implementation knowledge.
  • Software Engineers: the people who is expected to build the software and implement features.

The main goal: compiling software requirements

At the center is the overall activity called Compilation of Software Requirements. This is the process of collecting, writing down, checking, and organizing what the software should do.

Two groups can trigger and participate in that process:

  • DHwiki members
  • Digital Humanities practitioners

In other words: requirements are not written only by engineers; they are compiled collaboratively by the communities who know the goals and use cases.

What “compiling requirements” consists of

When the group compiles requirements, they always do three core things:

  1. Gather User Needs They collect real needs from users and stakeholders: problems, desired features, constraints, workflows, priorities, and context.
  2. Document Requirements They translate those needs into clear, written requirements so they can be understood, discussed, and implemented.
  3. Validate Requirements They check that requirements are correct and usable:
    • Do they match what people actually need?
    • Are they clear and unambiguous?
    • Are they feasible and consistent?
    • Can everyone agree this is what should be built?

Additionally, there is an optional step:

  1. Prioritize Requirements (sometimes) If needed, the group decides what matters most and what should be done first (because time and budget are limited).

Where the requirements are written and maintained

When the team documents requirements, they produce two concrete outputs:

  • A requirements spreadsheet This is a structured document in ODS spreadsheet format (a common open format). Think of it as the snapshot of requirements in a table, useful for overview, sorting, filtering, and reporting.
  • A set of tickets in the Wikimedia’s tracking system (Phabricator) Requirements are also turned into actionable, traceable work items using Wikimedia’s ticketing system, specifically under the DHwiki tag. Tickets are useful because they support discussion, assignment, progress tracking, and linking to code changes.

So, the process produces both:

  • a structured spreadsheet for a consolidated view, and
  • tickets for day-to-day workflow and implementation tracking.

How this connects to software development

There is also a separate activity: Software Development.

  • Software Engineers are responsible for software development.
  • Engineers also interact directly with the ticket system (they work from tickets, update them, and use them to manage development).

A key message is: engineers use the compiled requirements to decide what software features to implement. The requirements process feeds development.

The overall narrative in one paragraph

People from the DHwiki community, Digital Humanities practice, and Wikimedia development collaborate to compile software requirements. They do this by gathering user needs, documenting them clearly, and validating that they are correct—sometimes also prioritizing them. The documented requirements are maintained in two practical forms: a spreadsheet as a structured requirements document, and in the Wikimedia Phabricator[5] service as traceable tickets under the DHwiki tag. Software engineers then use those validated requirements from the tickets to build software features.

Structured requisites document

The spreadsheet is structured in rows for each requirements which are described in these columns:

  • ID (integer): unique requisite internal identifier. Also the Phabricator ticket number. Is relevant.
  • Feature (string): free text with a short name/title (often the human-friendly label).
  • Platform (enum): target environment/scope. Value should one from the set
    • WBS, applies exclusively to the Wikibase Suite.
    • WBStack, applies exclusively to the software artifact to operate a Wikibase SaaS, like Cloud.
    • WBS/WBStack, applies for both platforms.
    • Independent, independent of the software platform.
  • Description (free text): free text, narrative requirement statement or explanation.
  • Priority (enum): High | Medium | Low. The value is a subjective selection by the document editor.
  • Need (enum): to classify why this requirement is needed. Value is a subjective selection by the document editor and one from the set Deprecated | Future enhancement | Must-have | Should-have | Nice-to-have | Optional| Regulatory | Resolved | Should-have | Technical-debt    
  • Stability (enum): expected volatility/maturity of the current description. Value is a subjective selection by the document editor and one from the set  Stable | Moderate | Unstable | Finalized
  • Artifact type (enum): the kind of deliverable that would satisfy the requirement. Type can be a software product or an information resource, like an ontology. Value should one from the set:
    • Wikibase extension, a PHP component for the MediaWiki platform directly related with Wikibase.
    • Wikibase gadget, a javascript program, can be deployed in the system or in the user's personal configuration.
    • ontology, linked data ontology, probably expressed in OWL or RDF.
    • vocabulary, a controlled list of terms, probably expressed in SKOS or in a less structured system.
    • companion software, software applications outside the Wikibase/MediaWiki ecosystem which are required, recommended for a Wikibase implementation
    • bot, in the MediaWiki lingo refers to external scripts for automating tasks, usually run in a user computer.
  • Artifact name (string): concrete deliverable identifier/name (often “none at this moment” or a specific component name).
  • Source (reference string): traceability pointer, can be a person/org name, a URL, or an issue tracker reference (notably many rows use Wikimedia Phabricator ticket IDs like Txxxxx: ...).
  • Notes (free text): extra context.

Requisites

At this moment the requisites document is being curated in a Google Docs spreadsheet and is publicly accesible.

References

  1. OpenAI. (2025). ChatGPT (Dec 30 version) [Large language model]. https://chatgpt.com/
  2. Washizaki, H., & Olszewska, J. I. (Eds.). (2025). Guide to the Software Engineering Body of Knowledge v4.0a: Vol. 4.0a. IEEE Computer Society. https://ieeecs-media.computer.org/media/education/swebok/swebok-v4.pdf
  3. ISO Central Secretary. (2017). ISO/IEC/IEEE 24765: 2017(E): ISO/IEC/IEEE International Standard - Systems and software engineering--Vocabulary (p. 530) [Standard]. International Organization for Standardization. https://www.iso.org/standard/71952.html
  4. Rumbaugh, J., Jacobson, I., & Booch, G. (1999). The unified modeling language reference manual. Addison-Wesley.
  5. See https://phabricator.wikimedia.org/