The question behind the question
When a data platform comes up in the German Mittelstand, the question of staffing follows sooner or later. It is usually asked in one of two ways. One is whether this means hiring specialists who are barely available on the job market. The other is whether the existing IT team could not take it on as well.
The honest answer is closer to the first, at least if the result is meant to hold up over time and create value for years. A data platform is not software that you install and then administer. It is a built solution that connects a company’s systems, data, rules, and applications, and the skills required for that are spread across several disciplines.
A sound assessment has to distinguish between two phases and between two kinds of competence: depth of implementation and ownership of the subject matter.
Two phases with different needs
The build is limited in time and demanding while it lasts. Decisions are made that usually hold for a long time afterwards: the structure of the data models, how the source systems are connected, the permission model, the definition of the first metrics. Changing them later remains possible at any point, but it creates recurring effort, because the analyses and applications built on top have to be carried along. Getting it right from the start therefore pays off across the whole lifetime.
Operations run differently, but they are not the easier task. Well-designed operations are largely automated, and that is exactly where the effort sits: this automation has to be designed, calibrated, and continuously adapted to new sources, changed processes, and new requirements. On top of that, deviations do not get to choose which area they occur in. Whoever is responsible for operations has to be able to step in just as readily for a changed interface, an unclear metric, a permission issue, or the declining answer quality of an assistant. Over the year the time required is lower than during the build, the breadth of skills needed is the same, and it is needed at short notice and in full depth.
Which disciplines are involved
Seen from the outside, building a data platform is often described as one job. In fact several specializations are involved, and in practice they are rarely covered by the same people.
Data architecture and the semantic layer. Building the layers from raw delivery to an analyzable data product is one part. The more demanding one is the layer above it, where technical tables turn into a structure the business can understand: consistent terms, binding metric definitions, documented meaning for each field. This layer decides whether analyses say the same thing across the whole company and whether people and AI applications alike interpret the data correctly. Experience is least replaceable here.
Connecting the source systems. Every system delivers differently. Interfaces are documented to varying standards, change histories are reliable to varying degrees, data formats are clean to varying degrees. A robust connection requires knowing the typical failure patterns, and that knowledge builds mainly through the number of connections already made.
Platform and cloud infrastructure. Setting up and configuring the environment, automating the delivery of changes, separating test and production environments, availability, recovery, and controlling resource consumption. This field keeps moving quickly and requires continuous learning.
Security and permissions. Implementing a permission model down to row level, connecting it to the existing user directory, and making sure the rules take effect across every path of access. Mistakes in this area go unnoticed in normal operation and carry a correspondingly high cost.
AI. This field is considerably more demanding than the word integration suggests. It calls for specialist knowledge of which model suits which task, how data and its meaning are provided so that a model uses them correctly, how new model generations are integrated and tested before switching over, how answer quality is measured and improved systematically, and how to make sure an assistant uses only the data the person asking is entitled to. This discipline is changing faster than any other right now.
Application development. For the business teams to work with the platform, they need an interface that can be used without technical background, and someone to keep developing it. This part is regularly underestimated, yet it decides whether the data foundation that has been built is actually used day to day.
Operations. Monitoring the pipeline runs, quality checks, handling deviations, tracking changes, and a defined arrangement for cover.
On top of that comes a share that belongs to none of these technical disciplines and takes up considerable room in practice: gathering and sharpening requirements from the business teams, regular alignment with the people who will later work with the analyses, preparing decision papers on the points the company has to settle itself, such as priorities, metric definitions, and the business meaning of individual data fields, plus documenting the decisions taken and onboarding the users. This work is not a side effect. It largely determines whether what has been built technically lands in the company at all.
What this means for building a team in-house
This breadth raises three practical questions.
The first concerns staffing. Realistically this spectrum starts at two or three very experienced people, each of whom masters several of the areas named above at a high level and can stand in for the others. There is no ceiling, because as the number of connected systems, use cases, and users grows, so do the teams that handle this work in larger companies. What matters is less the headcount than the depth: every one of these areas is a specialization, and none can be built up on the side of another main job.
The second concerns utilization. During the build the demand is high; in the routine operation that follows it is lower, though it rises again with every system migration and every new requirement. Specialists on the payroll also have to be kept busy and developed further during the quieter phases.
The third concerns the job market. Experienced people in data architecture, platform operations, and AI applications are scarce, and most of them are tied up by larger companies and specialized service providers. And where a hire does succeed, it creates a dependency on individuals that is felt immediately when they fall ill or move on.
What has to come from the company
Set against this is a second kind of contribution that only the company itself can make, because it concerns its own business.
Ownership of definitions. Whether an order counts at the time it is placed or at the time it ships, which costs go into the contribution margin, and how cancellations are treated is not a technical question. The answer has to come from the company, and it has to hold there.
Knowledge of the upstream systems and processes. Every ERP or CRM that has grown over the years has fields repurposed for something else, historical special cases, and unspoken conventions. This knowledge sits with the key users in the business teams. Their availability is often the limiting factor during the build.
Decisions about access. Who is allowed to see which data is a management decision. The technical implementation is the platform’s job; the decision is not.
Prioritization. There are always more worthwhile analyses and use cases than can be built at the same time. The order they are tackled in largely determines how quickly a benefit becomes visible.
One qualification matters here, and it rarely appears in proposals: none of this has to be spelled out at the start. In very few companies are all metric definitions documented or all priorities agreed, and that is no reason to hold back. Working this out is a craft in itself, and there are methods for it: structured conversations with the business teams, decision papers with worked-out alternatives, documented decisions. Whoever supports the build should be able to guide that work. The input and the decision stay with the company; the work of getting there does not have to be done alone.
What AI adds on top
Adding AI changes the requirements on both sides.
On the implementation side, the discipline described above comes in, along with the need to keep it current. What counts as a good setup today can be outdated in twelve months. For a single company that is hard to keep up with; for specialists who work with it every day it is routine.
On the substantive side, clean definitions carry more weight. As long as analyses are produced by a handful of experts, those experts even out the fuzzy edges in their heads. Once questions can be asked freely, that stops happening. Access rules gain weight in the same way, because far more people ask far more questions. And a need for judgment arises in everyday use: employees have to be able to tell when a quick answer is enough and when the source should be checked. That is not a technical qualification but a matter of onboarding and internal rules.
Conclusion
Building and running a data platform with AI applications covers data architecture, system integration, cloud infrastructure, permissions, AI, and application development, plus operations that have to run reliably for years, and a substantial share of alignment, requirements work, and documentation.
Three questions follow from this for planning. Which of these skills does the company want to keep in-house for the long term, and can the utilization and development that requires be sustained? Where does it draw on experience that is applied every day somewhere else? And who inside the company owns the substantive decisions, with what mandate and with how much of their time?
How the answers turn out depends on size, system landscape, and the existing IT setup, and it differs from company to company. What matters is that the questions are answered before the start and not during implementation.