Beyond data: conversation with Zia Alborzi, Senior Data and AI Solution Consultant and Architect at NTT DATA

As Europe moves deeper into the age of AI, the central question is how data can be discovered, accessed and reused in a trusted, lawful and privacy-preserving way. In this interview, Zia Alborzi, Senior Data and AI Solution Consultant and Architect at NTT DATA, explains why data spaces are becoming essential infrastructure for Europe’s digital future — especially in health, AI, governance and cross-border collaboration.
Alexandru Dan from TVL Tech, a member of Ocean Enterprise Collective, is leading a series of conversations with experts across Europe to explore how data spaces are evolving and why they matter for the future of trusted data sharing, cybersecurity, AI and digital sovereignty. Through these interviews, he brings together practical perspectives from professionals working at the intersection of technology, governance and innovation.
Could you briefly introduce yourself and explain how you became involved in data space projects?
I’m Zia Alborzi. My background is in software engineering, data ecosystems, AI, high-performance computing, and architecture. I started my career quite deep on the technical side, building systems around data, computation, and distributed architecture.
Today, I work as a Senior Data and AI Solution Consultant and Architect, and I am also architect and technical lead for Dataspace4Health in Luxembourg. My route into data spaces was a bit of a winding one: I started with research, PhD work, high-performance computing, GPUs, software, and data systems, originally quite far from health. Over time, I moved from building individual systems toward building the architecture between systems: the layer that allows different organisations to collaborate around data without losing control, trust, or accountability.
In Dataspace4Health, the work is about taking European concepts such as Gaia-X trust, EDC-based exchange, verifiable credentials, federated catalogues, and secure processing environments, and turning them into running components for the health sector.
What pulled me in, though, came down to one realization: the old model just doesn’t fit sensitive data. You basically have two bad options. Either you say the data are locked in silos and we do nothing, and nothing ever happens. Or someone says, okay, let’s centralize everything in one big lake, and now you’ve got a legal, ethical, and security headache that nobody wants to own. And in health, you see this every single day.
For example, hospitals are sitting on incredibly valuable clinical data. A research group needs them for a study, and AI wants to validate the model. But nobody wants copies of uncontrolled and unverified patient data floating around. So the real question becomes: how do you make the data discoverable and reusable while it never actually leaves the control of the data holders or controllers responsible for that data?
That’s a puzzle that got me hooked into this data space, and it’s never about just one connector. It is the whole chain: creating the catalog, identity, credentials, permits, and so on and so forth, and the secure processing environment — in short, SPE. And for me, data spaces are the practical middle ground between two extremes that both fail.
Why do data spaces matter now in Europe, especially in relation to digital sovereignty and trustworthy data sharing?
To me, that matters now because three things have arrived at the same time in Europe: regulation, AI, and the whole conversation about sovereignty. And all three suddenly depend on being able to access good data in a governed way. All three.
On the regulation side, I can say Europe has done the hard part of setting the direction, putting the Data Governance Act, the AI Act, a lot of different acts as well, and EHDS specifically for health data. But here’s the thing: regulation on its own doesn’t make data flow; it is just a procedure.
We still need architecture, the operating model, and the trust services underneath to make sure this data exchange will happen. And that layer, I would say, doesn’t actually build itself. We have to build it.
And then AI is really forcing the issue, because everyone right now is talking about AI, and they want to create vertical AI, especially in health, energy, industry, and so on and so forth. But we cannot build serious AI in any sector that we can think of without good data, and the good data are fragmented, siloed, and locked away. You never have good AI without proper data.
I take health as an example because we are working on Dataspace4Health. EHDS is in force now, but the application is phased over the next few years.
And to me, that’s the transition window, which is a gift for us, in a way, because it’s a time to actually build the real capability: the access body, the catalog, the secure processing environment, cross-border access, audits, traceability, these kinds of things — all in the data space.
If we sleep on it, let’s say, if we don’t do anything, we arrive at 2029 and we have regulation, but we don’t have any proper architecture for the health domain.
And for the sovereignty part, I can say that I don’t read sovereignty as isolation or building walls, something that actually blocks everything. No, it’s not like that. For me, sovereignty is practical. It is the ability to define, enforce, and prove the rules under which data can be reused. A hospital should be able to say: this dataset can be used for this approved purpose, by this verified actor, in this secure processing environment, under this permit, with these output restrictions.
So sovereignty isn’t just saying that you shouldn’t share your data. It is being able to say yes under certain conditions that you can really enforce, not just talk about or promise.
What makes a data space different from traditional data platforms, data lakes or data marketplaces?
The real difference is that, in a data space, trust and control travel with the data. A lake usually centralizes the data, and a marketplace handles transactions, but a data space keeps a verified identity and usage rules attached to the asset.
Let me explain in a better way. A data lake is fantastic. It’s great inside one organization. But it assumes that there is just one owner, one governance model, and one boundary. The moment you’ve got many independent organizations with different legal duties or different interests, for instance, the model completely falls apart.
And the marketplace, for me, is very useful as well for discovery and payment. That’s great. But it usually stops at: here’s the dataset, here are the terms.
It doesn’t really solve the trust problems that we may face. A data space adds basically two things that neither of those has. For instance, it covers both the things that you have in the marketplace and in the lakehouse. And the two things that actually matter when multiple parties are involved are verifiable identity and machine-readable usage rules.
So what I can say to you is that, if you are exchanging right now in the data space, it isn’t just a file; it is an asset, plus the identity of who’s involved, plus the policy, plus a negotiated contract. It’s a whole package that you have in a data space.
And I can make this more concrete with the example of Dataspace4Health. We can onboard a participant. It’s not just creating a username and password to enter our platform or ecosystem. It’s issuing and checking verifiable credentials, and indeed a decentralized identifier and a certificate chain validating against the trust framework that we have.
And before anyone even talks about sensitive data, both sides can prove cryptographically who they are. I am a hospital; you are the research institute. That already puts it in a different universe from a normal platform, if you actually think from that angle.
Publishing a dataset is not simply uploading a file. The dataset has to be described as a data-space asset, with metadata that can be understood by both humans and machines.
In our case, this means mapping catalogue metadata, for example DCAT-AP-style descriptions or health-sector metadata profiles, into EDC assets. Then the access conditions are attached as policies, so when a request comes in, the connector negotiates against those policies before a single record moves.
I can say a marketplace also answers this: dataset availability. A data space can also answer much more than that, which is harder: who exactly you are and what you are allowed to do with this type of data or asset that you have. And we can both prove that the rules agreed on the usage of the data are respected. That’s a big difference between these three.
And then you also have traceability, auditability, and the transactions. So there are many steps in the process and clearly defined rules that are respected end-to-end.
Where do federated learning and compute-to-data fit into the future of data spaces?
To me, I see all of them as the execution layer of the privacy-preserving data space. It’s just the part where the work actually gets done without data moving.
And I would rank them by how quickly they pay off. Federated analytics gives you big value almost immediately, because a lot of the time people don’t need some huge AI model behind the scenes. They just need the answer to their simple question.
And compute-to-data is the next level. It’s an upgrade. It’s another layer for me, where an approved algorithm needs to run right next to the source of the data.
Federated learning is the most powerful for AI, basically, but it is not exactly plug-and-play. It still needs semantic alignment, privacy control, and model governance. It’s not just the magic that people think all the time about federated learning.
Let’s go a little bit into the research as well. For instance, if you want to do research between Luxembourg, France, and Germany, the first question probably is not about how to train a giant model. It’s much more about how many patients match this cohort, these biomarkers, these outcomes, this kind of thing. And federated analytics answers that, without ever pooling the raw data, for these kinds of questions.
The most important thing is that nobody’s data leaves the building, and that’s great. Then later, we might need to validate the algorithm. For instance, we have a very nice algorithm and we want to validate it; compute-to-data will come into the picture. The algorithm basically travels into the controlled environment, runs against the approved data — basically, it should be the real data — and only the permitted result comes back out.
The data stays under control, and the approved computation moves to the controlled environment. In a health-data-space context, this usually means running the computation inside a secure processing environment, with only permitted results leaving. And then, for training, federated learning lets each hospital train locally, basically, and share only the model updates.
And that’s great. Actually, this is an extension of the Dataspace4Health that we have, and we want to do it with Japan’s universities and hospitals. And I should say this caveat: those updates can still leak information — this is quite important — if the design and architecture are not very good, meaning that they are not bulletproof all the time. You have to take care of this.
And the model can also be biased if the underlying data are inconsistent, which is why, before any federated learning or federated AI, we need to actually do some boring stuff behind the scenes, which comes at the lower level.
And this boring stuff is metadata quality, common data models, agreed cohort definition, quality checks or data readiness, which is a very important and hot topic recently, and governance on the data.
So, federated analytics gives you questions at scale, compute-to-data gives you controlled execution, and federated learning gives you collaborative AI, but only if the governance and the semantics below that, in another layer, are actually mature enough.
What is the connection between AI, data spaces and data reuse governance?
To me, a data space is a kind of governed data supply chain for trustworthy AI. That’s the cleanest way I can put it.
Because AI needs data, but not just any data. It needs quality, lawful, and well-documented data. Otherwise, AI doesn’t work. And a data space is exactly the thing that can provide it with the paperwork attached, such as discovery, provenance of the data, and access conditions, which are always important for AI training, validation, and auditing as well.
These are all the things that let you trust the inputs, not just the model. So these are all related to data space.
And there’s one point that I feel strongly about: sovereign AI isn’t only about where the model runs. Everybody fixates on that. They think it’s the most important thing, but it’s just as important where the data comes from, under what rules, and for which AI model it can be used.
It is very important to know the lineage of those data as well. And that’s the second half that always gets forgotten: data. We are talking about AI all the time.
Having worked hands-on on both data and AI, I have also worked with sovereign AI infrastructure, including deploying open models such as Mistral 7B using vLLM on sovereign GPU infrastructure. And I would say that the most useful experience I got is that it showed the whole problem: running the model locally is not enough.
That’s only one part, and the part that people usually celebrate. But the harder questions are underneath. What domain of the data can this model actually use? Is it allowed for this purpose? Is it inside the secure processing environment? Can I trace the source? Can I prevent data leakage? Can I audit how the data is used in the training workflow?
In health data, this is not optional, because we are talking about sensitive data, and that’s the whole thing. A medical AI system does not earn trust simply because the model is powerful or has many billions of parameters. It earns trust through governed data reuse and the quality of the data that has been used for training those models.
So what I would like to sum up here is that, without governed data reuse, AI can be very powerful, but without governed data it’s fragile, it can fail in ways you can’t explain or defend.
So, for me, a data space is what makes the whole process for this AI accountable, because the data comes in a governed way and is used in AI.
Alexandru Dan: Yes, and with the AI Act, and also AI in health, we have a lot of regulations, and traceability, explainability, lineage, all these things have to be really well put in place. We cannot discuss it. It’s a rule that you need to obey. Because later, if you have bias, you need to know the source of the result. You need to trace it back to the data input.
What should Luxembourg focus on in the next 12 to 24 months to become more relevant in the European data space ecosystem?
My honest take is that Luxembourg shouldn’t try to win on size. It will never. We know Luxembourg is small, and that’s fine. It should win on focus, trust, speed, and cross-border interoperability. Take the actual strengths that we have in Luxembourg.
Concretely, in the health domain, the first priority should be to turn Dataspace4Health and HDAB-LU into real operational reference cases, not just strategic initiatives that look good in presentations. Second, we should connect health data spaces with sovereign AI infrastructure, because Luxembourg already has strong assets in high-performance computing and trusted digital infrastructure. Third, we should make the “boring plumbing” — onboarding, credentials, metadata, and access workflows — pragmatic and reusable, so each initiative does not reinvent the same foundations. Fourth, Luxembourg should stay aligned with European standards and initiatives such as EHDS, HealthData@EU, Gaia-X, DSSC, the Dataspace Protocol, and EDC.
And I’m very optimistic about Luxembourg specifically because it is small enough to coordinate, but connected enough to matter. That’s a rare combination.
In health, we could demonstrate the complete path end-to-end via Dataspace4Health: holder onboarding, cataloging, credentials, access governance, secure processing environment, and the output review that we do. Also, everything is traceable and auditable.
That’s a whole chain that is working in Dataspace4Health, and it can be used as a blueprint. And what I really want to say is that it is valuable for Europe because, right now, many countries are still translating EHDS from regulation into architecture, and Luxembourg has an opportunity to move faster by demonstrating a practical reference implementation.
If you look at it more generally, I’m not just talking about Dataspace4Health anymore; I am talking about HDAB-LU and the related national capabilities around health-data access, secure processing, and governance. It can become a kind of test case that everyone can look at. Then, of course, AI comes on top of it.
And if you can connect the privacy-preserving data space that I was talking about with sovereign compute, that’s a real niche, not just generic AI that is working. It should be trustworthy AI, governed through European data.
So the opportunity for me here is that we shouldn’t try to be big. We are small, but we can be credible and offer a reference implementation that others in Europe can learn from.
Alexandru Dan: I think credibility is really important these days with AI, data spaces, and also medical data. I think this is something that Luxembourg is really good at: establishing credibility. And then the technical and legal parts. I think Luxembourg is really good at respecting regulations, coming from the banking side. So I think we are well positioned for this, let’s say, more complicated projects where you need coordination, you need governance. And the good part is also the HPC and access to compute that we already have in Luxembourg.
What is the main message people should remember about data spaces?
To me, the promise of the data space is never about “here’s another platform.” Honestly, the last thing we need in Europe is another platform. The promise is a new model for collaboration.
And we should take into account that, once we’re talking about data space, we’re talking about a new model of collaboration. To me, Europe needs these data that are siloed for research, AI, healthcare, and industry. But we cannot unlock them by ignoring privacy, sovereignty, or even trust.
So we cannot actually unlock them by pretending that the answer is one giant central lake or data lake, or something where we can create a platform and use it in all European countries.
For me, data spaces are the architecture in between: not closed silos, and not uncontrolled centralisation. They allow governed, auditable, privacy-preserving reuse. That is the real promise: unlocking value from data while keeping control, accountability, and trust intact.

