Roark McColgan | Graph Massivizer EU Project https://graph-massivizer.eu Thu, 26 Feb 2026 06:34:00 +0000 en-US hourly 1 https://wordpress.org/?v=7.1 https://graph-massivizer.eu/wp-content/uploads/sites/27/2023/01/cropped-favicon-32x32.gif Roark McColgan | Graph Massivizer EU Project https://graph-massivizer.eu 32 32 DataNexus and EUDATA+: the clustering approach of Graph-Massivizer https://graph-massivizer.eu/datanexus-and-eudata-the-clustering-approach-of-graph-massivizer/ Mon, 16 Feb 2026 05:49:57 +0000 https://graph-massivizer.wp.itec.aau.at/?p=1424 With hundreds of projects funded by the European Commission that run more or less at the same time, activities that used to be rather easy in the past, have become a real nightmare for projects. Some of those activities are: i) understanding what other projects do and capitalize those developments and outcomes for your own work, ii) letting the wider audience know what is the positioning of your project in comparison to other projects working in similar fields -and thus, understanding complementarities, commonalities, differences, value proposition of each of them; so, using i) for the benefit of creating clear messages that help target communities understand the content-, and iii) attracting the attention of your target audience, meaning just being able to capture some minutes of researchers or policy makers in an era characterized by information overload.

In this complex context, the instrument of clusters emerges as a very useful tool not only to become effective in the aforementioned activities, but also to develop them in an efficient way.

Graph-Massivizer understood this premise from the beginning, and instead of setting up many new channels with a very limited reach, we created a strategy that would revolve around the principles of collaboration and networking. This has been materialized by the set up of two clusters that offer complementary opportunities to the project.

The first cluster, extremely critical to the success of Graph-Massivizer, has been labelled as DataNexus. It brings together all the Research and Innovation Actions addressing the topic of Extreme data mining, aggregation and analytics technologies and solutions. These projects aim to provide ground-breaking advances in the performance, speed and/or accuracy as well as usefulness of data discovery, collection, mining, filtering and processing when “extreme data” is involved.

Extreme data is defined here as “data that exhibits one or more of the following characteristics, to an extent that makes current technologies fail: increasing volume, speed, variety; complexity/diversity/multilingualism of data; the dispersed data sources; sparse/missing/insufficient data/extreme variations in values”.

According to IDC, the volume of data created each year is forecast to increase at a CAGR of 24.9% from 2024 to 2029 (faster unstructured data). In the case of Data Integration and Intelligence SW, the revenue at worldwide level is projected to nearly double from $6.4B in 2024 to $12.2B in 2029 (11,8% for EMEA), and interestingly enough, the AI Life-Cycle Software showcases CAGRs above 27% 2024-2029 (EMEA from $3B to $11B). Extreme data will influence all these new solutions, and has a huge potential market, as a great percentage of data falls under the former definition of extreme data. A lot of challenges come with those opportunities, such as complexity and integration, rising costs, regulatory and security concerns, talent and ecosystem gaps and ROI and value extraction, to name a few. Addressing these challenges requires a holistic approach and collaboration between those initiatives that focus on the different “pieces” of the big problems. Graph-Massivizer, for example, targets the whole lifecycle of graph-based data.

DataNexus, as a cluster that connects the different projects working on extreme scale data challenges has been instrumental in creating a knowledge base of technologies and developments, allowing project partners to collaborate in common aspects and understanding complementary views of addressing similar problems. In addition, the entire portfolio enables a more complete picture of the challenges that arise when dealing with the computing continuum, the processing of data in different computing infrastructures or issues associated to diverse vertical sectors. Furthermore, extreme data as a research topic has been highlighted by the strength of the cluster, which is more powerful than that of a single project. The following table summarizes key aspects about the positioning of the different projects and diversity of use cases covered, as well as contributions to the cluster.

DataNexus Cluster Overview

Project Objectives Focus Area Contribution to DataNexus
Graph-Massivizer – Extreme and Sustainable Graph Processing Graph-Massivizer develops methods and tools for extreme and sustainable graph processing to address urgent societal challenges that require extracting insights from complex relational data structures Digital twins for sustainable exascale computing
Green AI for automotive and industrial domains
Foresight modelling for environmental protection
Sustainable and green financial analytics
Graph-Massivizer’s expertise in scalable graph analytics enhances the cluster’s capacity to interpret complex inter-related data at scale, facilitating advanced analytics for use cases where relational structures are central
NEARDATA – Extreme Near-Data Processing Platform NEARDATA aims to build platforms that enable near-data processing, minimising data movement and enhancing responsiveness for extreme data workloads High-performance processing of genomics and metabolic data
Surgical data analysis and real-time insights
Novel architectures to support privacy and performance in sensitive data environments
By pushing computation closer to where data resides, NEARDATA addresses critical performance and privacy challenges inherent in extreme data analytics.
EXA4MIND – EXtreme Analytics for Mining Data Spaces EXA4MIND develops a platform for extreme data analytics, automation, and integration, particularly on HPC and supercomputing infrastructures. Automated data management integrated with European data ecosystems
Advanced analytics tools that support edge-to-HPC workflows
Analytics-as-a-Service (MAaaS) capabilities, e.g., for mobility risk forecasting and traffic flow analytics
EXA4MIND enhances the cluster’s capabilities in bridging high-performance computing with real-world analytics needs, particularly for mobility and large-scale event forecasting
EXTRACT – Distributed Data-Mining Platform EXTRACT focuses on distributed data-mining technologies that scale across heterogeneous infrastructures Personalised evacuation systems
Real-time distributed knowledge extraction
Cross-domain data mining for safety and resilience
EXTRACT brings scalable distributed mining capabilities, enabling the cluster to handle dynamic and geographically dispersed data sources
SYCLOPS – Cross-Architecture AI/Data Acceleration SYCLOPS is committed to democratising AI and data acceleration using open standards and cross-architecture solutions Hardware-agnostic acceleration frameworks
Standard-based AI/data toolchains
Accessibility and inclusivity in high-performance analytics
SYCLOPS strengthens the cluster’s technological foundation by lowering barriers to adopting accelerated computing across diverse hardware environments
EMERALDS – Extreme-scale Urban Mobility Data Analytics EMERALDS develops data-as-a-service and analytics platforms for urban mobility, emphasising scalability and privacy. Intelligent mobility analytics
Event risk assessment and forecasting
Integrated traffic management and flow analytics
Through real-world urban mobility use cases, EMERALDS grounds the cluster’s technologies in practical, impactful deployments that inform smart city development
EFRA – Extreme Food Risk Analytics EFRA targets risk analytics for food safety and supply chain resilience, leveraging extreme data to predict and manage risks Predictive models for food pathogens
Pest and contamination forecasting
Decision-support intelligence for regulatory frameworks
EFRA’s domain-specific analytics enrich the cluster’s multi-sector relevance, particularly for safeguarding food systems using advanced predictive insights.

The DataNexus cluster has produced a lot of materials that provide more elaborated insights of this work. See links below for reference [1].

EUDATA+ Cluster Overview

Project Objectives Focus Area Contribution to EUDATA+
Graph-Massivizer – Extreme and Sustainable Graph Processing Extreme data processing and analytics for complex data structures using massive graphs. Digital twins for sustainable exascale computing
Green AI for automotive and industrial domains
Foresight modelling for environmental protection
Sustainable and green financial analytics
Graph-Massivizer contributes scalable tools for data ingestion and analysis that help transform large, relational datasets into actionable knowledge pipelines — an essential component for data marketplaces and lifecycle orchestration within EUDATA+
PISTIS – Promoting and Incentivising Federated, Trusted, and Fair Sharing and Trading of Interoperable Data Assets Secure platform for sharing, trading, and monetizing proprietary data with technologies such as federated sharing and AI-driven quality assessment Mobility and Urban Planning
Energy
Automotive
PISTIS brings capabilities for trusted data exchange and trading infrastructure, underpinning monetization and governance solutions across cluster activities..
FAME – Federated decentralized trusted dAta Marketplace for Embedded finance Federated, trustworthy data marketplace facilitating monetization and trading of data assets, especially in the embedded finance domain with a strong emphasis on energy efficiency and security. Financial recommendation engine for families
Embedding Finance Services in a Personalized Citizen Wallet
Personalized Collaborative Intelligence for Enhancing EmFi Services
The EU Funds Application Process Made Easy
ESG Scorecard Ranking & Sustainable Portfolio Optimisation
Embedding Climatic Predictions in Property Insurance Products
Assessing the Quality and Monetary Value of Data Assets
FAME strengthens multi-sided data marketplace frameworks that interconnect producers and consumers across sectors, supporting sustainable monetization models
UPCAST – Universal Platform Components for Safe, Fair, Interoperable Data Exchange, Monetization and Trading. Tools and plugins to automate data-sharing agreements across multiple stakeholders, ensuring transparency and ease of use. Digital Marketing data and resources
Biomedical and genomic data sharing
Sharing Public Administration for climate across Thessaloniki cities
Health and fitness data sharing
Cactus marketing data
UPCAST brings practical workflow automation for contractual and technical data sharing, enabling seamless integration of distributed datasets in shared environments
enRichMyData – Empower AI-driven business products and services An open toolbox of scalable components for data enrichment, improving data quality, reusability, and value creation Marketing data Enrichment for smart-bidding optimization
Artificial Intelligence-based Welding Analytics
Service Data Enrichment for Smart Maintenance
European Register of Entities from Known Actions
Innovation Knowledge Graph for understanding Innovation lifecycle
Industrial Data Enrichment for Mineral Processing Optimization
By enhancing the quality and richness of data assets, enRichMyData supports the cluster’s mission to strengthen data value and enhance utility for analytics and monetization pathways.
DATAMITE – Monetization, Interoperability, Trading & Exchange Open-source framework to boost data monetization, interoperability, and exchange for diverse stakeholders including SMEs and public administrations Corporate Multi-Domain Data Exchange with DIH support
Corporate Multi-Site Data Exchange
Offering Data to Service Providers with DataSpaces
Leveraging Electricity Distribution Open Data
Connecting eDWIN to Data Markets
Connecting MISTRAL to the EU AI-ON-Demand Platform
DATAMITE contributes infrastructure and interoperability components for data sharing frameworks and exchange ecosystems, integral to cluster demonstrations and standards engagement
ExtremeXP – Experiment-driven and user-oriented analytics for extremely precise outcomes and decisions Human-centred analytics framework optimising complex data-driven workflows with integration of user preferences and feedback for personalised insights Crisis management
Cybersecurity
Public safety
Mobility
Manufacturing
ExtremeXP adds an experience-driven analytics dimension to cluster outputs, focusing on impactful, trustworthy insights derived from advanced data workflows

The EUDATA+ cluster has produced a lot of materials that provide more elaborated insights of this work. See links below for reference [2].

Conclusion

The set up of the DataNexus and EUDATA+ clusters and activities therein have been instrumental to give visibility to the outcomes generated by Graph-Massivizer to a wide audience. In addition to the increased number of dissemination opportunities (and thus Graph-Massivizer exposure), they have enabled a more clear positioning of our project in a complex ecosystem of projects that work in related fields, allowing us to derive concrete messages to our target audiences and to define a more accurate and finetuned value proposition, both aspects of great importance to foster the adoption of project results, which is one of the ultimate goals of the committed investments.

Author: Nuria de Lama (Consulting Director, IDC)

References

[1] https://www.youtube.com/watch?v=CLBs7Si0MNo;
https://www.youtube.com/watch?v=7kTdhwvELB4&t=139s
https://extract-project.eu/introducing-datanexus/
https://emeralds-horizon.eu/synergies/data-nexus-cluster

[2] https://datamite-horizon.eu/eudata/
Working Groups of the EUDATA+ cluster (link to zenodo)

]]>
GraphMa: Public Release and What’s New https://graph-massivizer.eu/graphma-public-release-and-whats-new/ Wed, 04 Feb 2026 14:54:39 +0000 https://graph-massivizer.wp.itec.aau.at/?p=1410 One year ago, we published a blog post introducing the concept of GraphMa within the Graph-Massivizer project. At that time, the GitHub repository was private, and we shared only conceptual insights and early design principles. During 2025, we made the GraphMa GitHub repository public.

GraphMa is now public, fully documented, and ready for developers to explore. You can now explore, clone, and contribute to GraphMa on GitHub:

The repository includes:

  • Core implementation of GraphMa’s pipeline-oriented graph processing model.
  • Benchmark suites for ingestion and traversal performance.
  • Examples and starter pipelines to help developers get up and running quickly.

GraphMa is fully open source under the Apache License, Version 2.0.

What Makes GraphMa Different?

GraphMa introduces four key innovations for large-scale graph processing:

  1. Pipeline Representation – Centralised construction and execution with type safety, defined as “blueprints” and evaluated lazily, enabling optimisations such as operator switching and resource-aware execution.
  2. Operator Model – Provides an extensible catalogue of graph analytics (e.g., centrality, clustering, structural metrics), including stateless transformations, stateful operations, and terminal triggers, supporting both generic and domain-specific logic.
  3. Directed Data Transfer – Implements a protocol for clear producer-consumer roles and deterministic message flow across pipeline stages.
  4. Higher-Order Traversal – Abstracts traversal mechanics into a unified protocol with modes for fine-grained, bulk, and conditional iteration, supporting scalable processing across diverse graph formats and large datasets.

These innovations enable high throughput and low latency for graph ingestion and traversal.

Getting Started

To start using GraphMa:

  1. Clone the repository: git clone https://github.com/graphmassivizer/graph-inceptor-graphma
  2. Follow the setup guide in the README.
  3. Explore the sample pipelines and benchmarks.
]]>
Navigating the Computing Continuum: Enabling Scalable and Sustainable Graph Processing https://graph-massivizer.eu/navigating-the-computing-continuum-enabling-scalable-and-sustainable-graph-processing/ Tue, 03 Feb 2026 08:40:57 +0000 https://graph-massivizer.wp.itec.aau.at/?p=1405 Rethinking Infrastructure Boundaries for Graph Analytics

The Graph-Massivizer project reimagines how large-scale graph data is processed by leveraging the computing continuum, a seamless integration of edge, cloud, and high-performance computing (HPC) environments. Unlike traditional siloed infrastructures, the continuum enables cross-layer orchestration, where graph workloads dynamically shift between infrastructure tiers based on real-time latency, energy, and cost trade-offs. This shift is crucial for modern data-intensive applications, where graph-based workloads, ranging from social network analysis to anomaly detection and recommendation systems, demand high throughput, low latency, and adaptive execution environments.

Why Graph Workloads Need the Continuum

Graph analytics pose unique challenges: irregular computation patterns, data dependencies, and unpredictable workloads. A single computation can touch millions of nodes and edges, triggering cascading execution chains. The continuum provides the flexibility needed to manage this complexity:

  • Latency-aware execution: Low-latency components like subgraph filtering or anomaly inference run directly on edge nodes (e.g., Jetson, Raspberry Pi), minimizing data transfer.
  • Sustainability-driven scheduling: Carbon-intensive operations, such as matrix multiplications or full-graph traversals, are routed to green cloud regions or HPC clusters with monitored energy efficiency.
  • Cost-efficient scalability: By adopting serverless computing, resources scale elastically with demand, avoiding idle infrastructure and reducing operational overhead.

Graph-Choreographer: The Serverless Orchestrator of the Continuum

At the heart of this adaptive orchestration lies Graph-Choreographer, the execution backbone of the Graph-Massivizer toolkit. It bridges graph workflows, known as basic graph operations (BGOs), with serverless, Kubernetes-native deployment across heterogeneous nodes.

Key Features

  • Declarative to Executable: Transforms high-level BGO workflows into executable DAGs using orchestration services like HEFTLess and EnergyLess.
  • Backend Agnosticism: Dynamically selects between Argo Workflows and OpenFaaS, enabling stateful or stateless execution depending on runtime conditions.
  • Real-time Adaptation: Monitors energy draw, CO2 intensity, and performance KPIs via Prometheus, Kepler, and PowerJoular, enabling runtime adaptations like reassigning tasks, throttling concurrency, or reprioritizing stages.


Graph Processing Continuum Cycle

 

The Next Frontier

The computing continuum challenges us to rethink the boundaries of infrastructure, not as isolated layers, but as interconnected zones of opportunity. For graph processing, this means more than just faster analytics; it means smarter deployments, energy-aware execution, and systems that respond to both workload and world conditions. As data grows in scale and complexity, our tools must evolve to be not only scalable and sustainable but also intelligent. Embracing the continuum is not just a technical decision; it is a strategic shift toward architectures that are adaptive by design and responsible by default. By uniting serverless computing, edge intelligence, and green HPC under a common orchestration model, projects like Graph-Massivizer pave the way for a new generation of data-driven systems that are both high-performing and environmentally conscious. The path forward is clear: if our data spans the continuum, so must our computation.

Author:
Dr. Reza Farahani
University of Klagenfurt, Austria

]]>
Synthetic Financial Data Generation: Engineering Market-Consistent Time Series for Quantitative Research, Trading, and Regulatory Compliance https://graph-massivizer.eu/synthetic-financial-data-generation-engineering-market-consistent-time-series-for-quantitative-research-trading-and-regulatory-compliance/ Tue, 27 Jan 2026 05:53:38 +0000 https://graph-massivizer.wp.itec.aau.at/?p=1397 Modern quantitative finance is increasingly constrained not by a lack of ideas, but by limitations in data availability, usability, and regulatory permissibility. As trading strategies, risk engines, and AI-driven models become more sophisticated, the demand for large-scale, high-fidelity financial datasets has gone beyond what real historical data can sustainably provide.

Synthetic financial data generation is emerging as a core capability rather than an experimental add-on. When engineered correctly, synthetic data enables quantitative teams to scale research, stress models beyond observed regimes, train AI systems robustly, and satisfy regulatory and compliance requirements, all without compromising market realism.

This blog, reflecting the work done and concluded in the Graph-Massivizer EU project, outlines how market-consistent synthetic time series can be engineered and generated using a graph-centric approach, and why this methodology represents a step change for quantitative research, trading, and regulatory validation.

Why Real Market Data Is No Longer Sufficient

While historical market data remains indispensable, it exhibits several structural limitations:

Limitation Details
1 Finite coverage of regimes Rare events (liquidity crises, volatility explosions, regime shifts) are under-represented or entirely absent.
2 Sampling and survivorship biases Many datasets are filtered, adjusted or incomplete, especially across long horizons.
3 Restricted scalability High-resolution data (minutely, tick or sub-second) becomes prohibitively expensive and operationally heavy at scale.
4 Regulatory and licensing constraints Reuse, redistribution, and model training are often limited by vendor agreements and compliance rules.
5 AI model brittleness Machine learning systems trained on narrow historical distributions tend to overfit observed regimes and fail under stress.

Synthetic data, when naively generated, risks compounding these problems. When engineered with market structure awareness, however, it becomes a strategic asset.

Defining “Market-Consistent” Synthetic Financial Data

Market-consistent synthetic data is not defined by point-wise similarity to historical prices, but by preservation of structural, statistical, and relational properties that govern real markets.

Key consistency dimensions Details
1 Statistical fidelity Distributional properties of returns, volatility clustering, heavy tails, skewness, kurtosis, and autocorrelation structures
2 Temporal dynamics Multi-scale dependencies across intraday, daily, and longer horizons.
3 Cross-asset relationships Correlations, co-movements, lead-lag effects, and regime dependencies.
4 Market microstructure constraints Plausible price formation, liquidity effects, and volatility-volume interactions.
5 Regime coherence Stability of relationships within regimes and realistic transitions between regimes, namely preservation of stable statistical and relational structures within a market regime.

Achieving these properties simultaneously requires moving beyond purely parametric models or black-box generative AI.

Graph-Massivizer Approach: A Graph-Centric Paradigm for Synthetic Data Engineering

Graph-Massivizer financial use case was built on the premise that financial markets are naturally relational systems, not independent time series collections. Assets, time steps, market regimes, and derived features form a structured network of dependencies that can be explicitly modeled.

Historical financial data across assets, instruments, and time resolutions is first ingested and transformed into a graph representation:

  • Nodes can represent time points, instruments, regimes, or derived states.
  • Edges can encode temporal transitions, cross-asset dependencies, and statistical constraints.
  • Multi-layer graphs capture interactions across different time scales.

This representation preserves information that is typically lost in flat tabular datasets. Then, before any generation occurs, the source data undergoes structural analysis to define what must be preserved and where variability is allowed. Next, synthetic data is then generated by expanding the graph under explicit constraints such as local randomness, correlations preservation where specific behaviors may be amplified, as well as reduced similarity to the original historic data to the point where reverse engineering is not possible.

The objective is original plausible novelty: data that is statistically consistent yet not traceable to any original observation, given that a critical compliance requirement is that synthetic data must not allow reconstruction of original data. This is particularly relevant for regulatory audits and third-party model validation.

Applications in Quantitative Research and Trading

Strategy Research and Backtesting Synthetic datasets allow quantitative teams to:
Alternative history Run thousands of alternative histories for the same strategy
Regime changes Evaluate sensitivity to regime changes and tail events
Over-fitting Reduce false confidence driven by over-fitted historical periods
Strategy Robustness Test strategy robustness under unseen market conditions

Performance metrics derived from synthetic data are diagnostic, not predictive, highlighting fragility and structural bias.

AI and Machine Learning Training For AI-driven trading systems, synthetic data provides:
Training Massive, balanced training corpora across regimes
Reduced Overfitting Reduced overfitting to dominant historical patterns
Generalization Improved generalization under volatility shifts
Compliance Safe experimentation without breaching data licenses

Synthetic data is a pre-training and stress-training substrate and not a replacement for real data.

Model Validation and Risk Stressing Risk and model validation teams can leverage synthetic data to:
Stress scenarios Generate extreme but coherent stress scenarios
Validation Validate model behavior outside observed history
Perturbations Compare model responses across controlled perturbations
Robustness Document robustness in regulatory submissions

This shifts validation from retrospective justification to proactive resilience testing.

Regulatory and Compliance Advantages From a regulatory standpoint, market-consistent synthetic data addresses multiple concerns simultaneously:
Data lineage and licensing Synthetic datasets can be shared internally and externally (with appropriate derived works redistribution license from historic data providers).
Model risk management Regulators increasingly expect evidence that models behave sensibly outside calibration samples.
Auditability Graph-based generation pipelines are deterministic, inspectable, and reproducible.
Privacy and confidentiality While financial market data is not personal data, irreversibility remains essential for proprietary and contractual protection.

Synthetic data becomes a compliance enabler

There are various challenges as not all synthetic financial data is fit for purpose. Common failure modes include over-fitting synthetic data to historical distributions, ignoring cross-asset and temporal dependencies, excessively smooth or overly random time series or lack of quantitative validation metrics. A graph-centric, constraint-driven approach mitigates these risks by design.

Conclusion

Synthetic financial data generation is no longer an experimental direction. When engineered with structural awareness, statistical rigor, and regulatory foresight, it becomes a core infrastructure capability for modern quantitative organizations.

Graph-Massivizer demonstrates that markets can be expanded, not merely replayed, producing market-consistent time series in extreme volumes that support deeper research, more resilient trading systems, and more credible regulatory validation.

The future of quantitative finance will belong not only to those who analyze history best, but to those who can systematically explore what history did not contain, while following the rules that markets, mathematics and regulators impose.

Laurentiu Vasiliu, founder, Peracton Ltd

19/12/2025

]]>
From Scarcity to Scale: How Synthetic Financial Data Is Powering AI Training for Quantitative Trading Strategies https://graph-massivizer.eu/from-scarcity-to-scale-how-synthetic-financial-data-is-powering-ai-training-for-quantitative-trading-strategies/ Mon, 12 Jan 2026 11:39:20 +0000 https://graph-massivizer.wp.itec.aau.at/?p=1393 Artificial intelligence has moved from experimentation to production across quantitative trading, portfolio construction, execution optimization, and risk management. However, while model architectures and compute capacity have scaled rapidly, the availability of high-quality financial data has not kept pace.

Buy-side quantitative teams face a structural constraint: financial market data remains scarce, fragmented, expensive, and often unsuitable for large-scale AI training. Historical datasets are finite, heavily reused, biased by survivorship and regime persistence, and increasingly subject to restrictive licensing terms. As a result, many AI-driven trading initiatives stall not due to lack of modeling sophistication, but due to insufficient, contaminated, or non-scalable data.

Synthetic financial data is emerging as a strategic solution to this bottleneck, enabling a transition from data scarcity to data scale, while preserving market realism and regulatory relevance.

 

Why Traditional Market Data No Longer Scales for AI Training

 

Quantitative trading strategies based on machine learning and deep learning differ fundamentally from traditional statistical or factor-based models. They require large volumes of diverse training data, exposure to multiple market regimes, including rare and extreme events, a clean separation between training, validation, and stress-testing datasets and continuous refresh without historical leakage or overfitting.

Traditional market data is limited on several of these dimensions:

Dimension Details
1 Finite history Even the most liquid instruments offer only a limited number of statistically independent samples once regime clustering and autocorrelation are considered.
2 Hidden data contamination Widely reused historical datasets introduce indirect information leakage across research teams, vendors, and models.
3 Cost and licensing constraints Scaling from gigabytes to terabytes of tick-level data is often economically prohibitive, particularly for smaller or mid-size buy-side firms.
4 Poor coverage of tail events Extreme scenarios such as flash crashes, liquidity gaps, structural breaks are precisely what AI models need to learn, yet they are underrepresented in historical data.

 

These constraints are structural, not incremental. They cannot be solved by marginally better data sourcing or vendor negotiation.

 

Synthetic Financial Data: From Approximation to Market-Consistent Engineering

 

Modern synthetic financial data is not a simplistic resampling or noise-augmented replica of historical prices. When engineered correctly, it represents a market-consistent multiverse of financial time series that preserves statistical properties across time scales, cross-asset and cross-market dependencies, microstructure dynamics (order flow, spreads, volatility clustering), regime transitions and structural breaks This can be achieved through a combination of stochastic and regime-switching models, graph-based dependency modeling and constraint-driven generation aligned with real market invariants

The result is not one synthetic dataset, but thousands—or millions—of plausible market trajectories that extend far beyond what history alone can provide.

 

Powering AI Training at Scale

 

Synthetic financial data fundamentally changes how AI models are trained and validated in quantitative trading.

Features Details
1 Unlimited data AI models benefit from exposure to orders of magnitude more data than historical markets can supply. Synthetic generation enables:Unlimited time series length
Massive scenario expansion
Parallel simulation across assets, venues, and regimesThis extreme data volume supports more robust representation learning and significantly reduces overfitting.
2 Controlled Regime Coverage Synthetic data allows explicit control over market regimes, including:
High-volatility and crisis environments
Illiquid and fragmented markets
Structural transitions (policy shifts, market microstructure changes)
Models can be trained not just on “what happened,” but on “what could plausibly happen.”
3 Clean Model Validation and Stress Testing By construction, synthetic datasets can be strictly partitioned, eliminating implicit look-ahead bias. This enables:
Cleaner backtesting
More reliable out-of-sample validation
Scenario-based stress testing aligned with regulatory expectations

 

Business Impact for Buy-Side Quantitative Teams

 

From a business perspective, the adoption of synthetic financial data is less about experimentation and more about competitive positioning. Quant teams can iterate models faster without waiting for new historical data or negotiating incremental licenses. Then, synthetic data decouples AI scaling from data vendor pricing, enabling predictable and controllable cost structures. Further on, exposure to a broader market multiverse improves resilience across regimes, directly impacting drawdown control and long-term performance stability.

Synthetic datasets support explainability, reproducibility, and scenario-based validation that are key concerns for internal model risk committees and external regulators.

What is changing today is not just the quality of synthetic financial data, but its role in the quantitative stack. It is evolving from an augmentation tool into core data infrastructure for AI-driven trading.

 

Scaling Artificial Intelligence, Not Just Data

 

Synthetic financial data enables this shift from scarcity to scale, by providing the foundation required for industrial-grade AI training in finance. For buy-side quantitative teams, it represents not only a technical advancement, but a strategic lever: accelerating innovation while improving robustness, compliance readiness, and long-term performance sustainability.

The evolution of quantitative trading will not be determined solely by better models or faster hardware, but by the ability to systematically train AI across diverse, realistic, and unbiased market environments.

In an environment where alpha is increasingly driven by adaptability rather than historical coincidence, synthetic financial data is rapidly becoming a requirement rather than an option.

 

Laurentiu Vasiliu, founder, Peracton Ltd

26/12/2025

]]>
How Temporal Shifting Affects the Carbon Intensity of Data-Centre Workloads (VU) https://graph-massivizer.eu/how-temporal-shifting-affects-the-carbon-intensity-of-data-centre-workloads-vu/ Mon, 22 Dec 2025 10:55:38 +0000 https://graph-massivizer.wp.itec.aau.at/?p=1386 Processing large-scale graph-processing workloads requires similarly large-scale infrastructure, which we know today as data centres: large computing facilities, deploying hundreds or thousands of interconnected computers. Data centres form the basis of today’s digital infrastructure and are necessary for a wide range of societally important tasks, ranging from facilitating government tax administration to sharing social media posts. By combining the capabilities of many machines, the computers in a data centre can complete computationally intensive tasks such as massive graph-processing workloads.

However, powering such large numbers of computers takes a significant and growing amount of energy. The global energy demand of data centres is estimated to reach 8% in 2030 [1]. Unfortunately, burning fossil fuels remains an important and widely used source of energy. This makes data centre construction and operation a significant contributor to greenhouse gas (GHG) emissions.

In this post, we work towards sustainable massive graph-processing workloads by using Graph Greenifier to analyse the effect of temporal shifting (a technique to run workloads with sustainability in mind) on the carbon emission of data centres.

 

What is Carbon Intensity?

 

Carbon intensity is a metric that quantifies the (un)sustainability of an energy source by computing the amount of CO2 emitted per unit of energy. The table below presents an overview of the carbon intensity of four highly popular energy sources [1].

 

Source Carbon Intensity (CO2/kWh-eq)
Wind 11
Solar 41
Oil 650
Coal 820

 

Today, in practice, the electricity that is generated from both renewable and non-renewable sources is combined and offered to electricity consumers through an electricity network, also known as the energy grid. This means that electricity consumers on the grid do not use entirely renewable or non-renewable electricity, but rather use whatever combination is put onto the grid by electricity producers.

By knowing how much electricity on an energy grid comes from each source, we can compute the carbon intensity of that grid. To do so, we use the following formula:

 

Formula for computing grid carbon intensity.

 

Here, CIg is the carbon intensity of the grid, S is the collection of energy sources used on the grid, CIs is the carbon intensity of one specific energy source, Es is the amount of energy obtained from that energy source, and Eg is the total amount of energy on the grid. In plain English, this formula computes the carbon intensity of each energy source and then computes the weighted sum of those sources.

Knowing the carbon intensity of a grid allows electricity consumers, such as data centres running massive graph-processing workloads, to compute the carbon emissions of their activities by multiplying their electricity use by the carbon intensity of the grid they use to obtain their electricity. Once the carbon emission of a data-centre workload is known, we can start exploring approaches such as temporal shifting to reduce it.

 

What is Temporal Shifting?

 

The carbon intensity of grids that obtain energy from renewable sources can change significantly over time because the amount of renewable energy is highly variable and depends on factors such as the time of day and the weather. For example, the image below shows the Dutch energy grid over the course of one month. The top plot shows the amount of available renewable (green) and non-renewable (gray) energy on the grid, and the bottom plot shows the carbon intensity of the grid.

 

Energy mix and carbon intensity of the grid in the Netherlands during October 2023 [6].

 

Intuitively, we can reduce GHG emissions by using electricity when its carbon intensity is low. We can implement this approach for graph processing in data centres by delaying the execution of incoming workloads when carbon intensity is high and starting execution when carbon intensity is low. This effectively moves the workload in time, which we call temporal shifting. This idea has been suggested in related scientific work [3, 4, 5], but Graph Greenifier allows us to simulate, and therefore quantify, its effect.

 

The Graph Greenifier Approach

 

Graph Greenifier can simulate what happens in data centres during massive (graph-)processing workloads and supports simulating operational techniques such as temporal shifting. Operational techniques are actions data centres take to influence their operation. For example, selecting when and where to schedule a task. This allows data center operators, designers, and researchers to understand the impact of such techniques on both data center performance and sustainability. The figure below shows a simulation result from Graph Greenifier for this scenario.

 

Simulation results comparing FCFS scheduling and Carbon-Aware scheduling.

 

Massive graph-processing workloads and other large workloads consist of many small tasks that need to be executed. The top plot in the figure shows the number of actively running tasks for two different scheduling approaches. The blue curve shows a traditional First-Come-First-Served (FCFS) scheduler, which schedules tasks to be executed as soon as possible, and in the order they arrive. The orange curve shows the Carbon-Aware scheduler, which delays the execution of incoming tasks when carbon intensity is high. The green curve in the bottom plot shows the carbon intensity of the grid over time.

We can see that the Carbon-Aware scheduler effectively delays tasks until carbon intensity is low by looking at the orange and green curves. Specifically, we see that the peaks in the orange curve (high number of active tasks) align with the valleys in the green curve (low carbon intensity). For this particular workload, the reduction in carbon emissions is 2.5%, but this can increase depending on the workload and the carbon intensity (variation) of the energy grid to which the data centre is connected.

 

Next Steps for Graph Greenifier

 

Temporal shifting is but one technique in a large collection of commonly used operational techniques in data centres. These include spatial shifting, checkpointing, and active-active replication, to name but a few. Additionally, changing the data-centre scheduling policy and other operational techniques can affect not only carbon emissions, but also workload performance and other non-functional properties.

By supporting these techniques in Graph Greenifier, scientists, data centre operators, and other stakeholders can explore “what-if” scenarios and perform a wide range of deep analyses using arbitrary combinations of these techniques to make a trade-off between sustainability and performance for their graph-processing workloads.

 

References

 

[1] Anders S. G. Andrae and Tomas Edler. 2015. On Global Electricity Usage of Communication Technology: Trends to 2030. Challenges 6, 1 (2015), 117–157. LINK

[2] Udit Gupta, Mariam Elgamal, Gage Hills, Gu-Yeon Wei, Hsien-Hsin S. Lee, David
Brooks, and Carole-Jean Wu. 2022. ACT: designing sustainable computer systems with an architectural carbon modeling tool. In Proceedings of the 49th Annual
International Symposium on Computer Architecture (New York, New York) (ISCA
’22). Association for Computing Machinery, New York, NY, USA, 784–799. LINK

[3] T. Sukprasert, A. Souza, N. Bashir, D. Irwin, and P. Shenoy, “On the limitations of carbon-aware temporal and spatial workload shifting in the cloud,” in EuroSys, 2024.

[4] Philipp Wiesner, Ilja Behnke, Dominik Scheinert, Kordian Gontarska, and Lauritz Thamsen. 2021. Let’s wait awhile: How temporal workload shifting can reduce carbon  missions in the cloud. In Proceedings of the 22nd International Middleware Conference. 260–272.

[5] Jiechao Gao, Haoyu Wang, and Haiying Shen. 2020. Smartly handling renewable energy instability in supporting a cloud datacenter. In 2020 IEEE international parallel and distributed processing symposium (IPDPS). IEEE, 769–778

[6] D. Niewenhuis, S. Talluri, A. Iosup, and T. De Matteis, “Footprinter: Quantifying data center carbon footprint,” in HotCarbon, 2024.

]]>
From Complex Queries to Intelligent Assistants: How Graph-Powered AI is Transforming Data Center Operations (UNIBO) https://graph-massivizer.eu/from-complex-queries-to-intelligent-assistants-how-graph-powered-ai-is-transforming-data-center-operations-unibo/ Thu, 18 Dec 2025 14:42:25 +0000 https://graph-massivizer.wp.itec.aau.at/?p=1377 Innovations from the University of Bologna and CINECA in the Graph-Massivizer Project

Authors: Prof.Andrea Bartolini (Associate Professor at University of Bologna) and Junaid Ahmed Khan (PhD student and Research Fellow at the University of Bologna)

Modern data centers and high-performance computing (HPC) systems generate extraordinary volumes of telemetry data. The CINECA Marconi100 supercomputer alone produced approximately 49 terabytes of uncompressed operational data, with individual systems containing up to one million unique sensors sampling at 20-second intervals [1]. This data flows continuously from compute nodes, memory subsystems, power delivery infrastructure, cooling systems, and storage components—creating an unprecedented wealth of operational intelligence. Yet despite this abundance, extracting actionable insights remains remarkably difficult.

The fundamental challenge lies not in data collection but in data access. Operators seeking to understand system behavior must navigate three simultaneous barriers: deep domain expertise in HPC operations, intimate knowledge of the specific monitoring framework architecture, and proficiency in the query languages and APIs of underlying NoSQL databases. This triple requirement effectively restricts analytical capabilities to a small group of specialists, leaving the vast potential of operational telemetry largely untapped.

Researchers at the University of Bologna, in collaboration with CINECA (Italy’s largest supercomputing center), have spent several years addressing this challenge through the EU-funded Graph-Massivizer project. Their work traces an evolution from complex manual queries through structured knowledge representations to an intelligent natural language assistant capable of answering operational questions with over 92% accuracy—transforming how facility managers and engineers interact with their data.


Figure 1: Examon’s massive scale and data heterogeneity

The Marconi100 (M100) system at CINECA utilizes a holistic monitoring framework for operational data analytics called “ExaMon”. It is designed to collect data from various sources, including hardware sensors, software logs, and performance metrics, and stores this data in a NoSQL database (Cassandra, with KairosDB for time-series) in a centralized repository. Figure 1 shows the complexity of the data collected at the M100 system using the ExaMon framework. It integrates data from nine specialized plugins that collect a wide spectrum of information—from air-conditioning and power-distribution data (Vertiv, Schneider, Logics) to node-level sensor telemetry (IPMI), cluster-wide performance metrics (Ganglia, SLURM, Nagios), and external environmental conditions (Weather). Each plugin contributes its own set of metrics and plugin-specific fields, resulting in a highly heterogeneous and multidimensional dataset. This diversity of data sources underscores the complexity of the monitoring environment.

Structured Understanding Through Domain Ontology

The research began by recognizing that the schema-less nature of NoSQL databases—while enabling flexibility and scalability—creates fundamental obstacles for complex analytical queries. Without predefined schemas, users must manually establish connections between different data sources, navigate vendor-specific naming conventions, and construct multi-step query chains that require extensive domain knowledge.

The team’s first innovation was developing a domain-specific ontology for operational data analytics telemetry [2]. Unlike existing data center ontologies that focus primarily on inventory cataloging and infrastructure documentation, this ontology captures the critical relationships between topological components, hardware systems, and job execution data. The representation mirrors how domain experts actually conceptualize HPC operations: racks contain compute nodes positioned in three-dimensional space, plugins organize sensors that produce timestamped readings, and submitted jobs connect directly to the computational resources they utilize. Figure 2 shows the developed ODA ontology.

Figure 2: Operational data analytics (ODA) Ontology

This structured representation enables queries through SPARQL, a graph query language whose logical structure follows the natural relationships in the data. Comparative analysis revealed dramatic simplifications: complex queries requiring dozens of lines of Python code with multiple sub-queries in NoSQL approaches could be expressed in just five to fifteen lines of SPARQL. More significantly, these queries follow paths that non-experts can understand by tracing relationships through the ontology—from rack to node to plugin to sensor to reading—without requiring intimate knowledge of database internals.

The approach proved particularly valuable for queries involving relationships across data sources. Calculating the average power consumption during a specific job’s execution, for instance, traditionally requires querying the job table to identify execution times and allocated nodes, then separately querying sensor tables for power readings during those intervals, and finally correlating the results through manual data manipulation. With the knowledge graph approach, this becomes a single query that traverses the explicit job-to-node-to-sensor-to-reading pathway defined in the ontology.

Confronting the Scalability Challenge

While the ontology approach proved effective for query simplification, a fundamental scalability challenge emerged during implementation. Converting time-series telemetry data into RDF triples resulted in storage requirements approximately 745 times larger than equivalent NoSQL representations. For a single month of data from just one monitoring plugin on the Marconi100 system, this translated to nearly three terabytes of graph storage versus four gigabytes in compressed Parquet format. Such overhead rendered full materialization impractical for production environments.

The solution emerged through virtual knowledge graphs—a technique that constructs graph representations on-demand, containing only the data relevant to answer a specific user query. Rather than materializing the entire dataset as a persistent graph, the system dynamically extracts entities from user questions, fetches relevant data from the NoSQL datalake based on identified time ranges and metrics, constructs a temporary lightweight graph, and executes queries against this focused representation.

Figure 3: Data Analytics (DA) Chatbot

Figure 3 shows the proposed end-to-end architecture for the Data Analytics (DA) chatbot for data centers and HPC [3]. The architecture contains five components: (1) Frontend, built using streamlit library, that provides a chatbot interface to the user, (2) Backend, that is built using FlaskAPI, once it receives a user question, two concurrent processes of SPARQL generation and Virtual Knowledge Graph (VKG) generation is started. SPARQL is generated using the (3) LLM inference service, that takes as input the user input prompt in natural language along with the ODA ontology, to guide the LLM about the logical data model, and a set of few-shot examples as context. Meanwhile, the second concurrent task of VKG takes the user prompt, and using natural language processing tools, such as pattern matching and rule-based critertias to extract the necessary entities from the user text that guides the system of which parameters and sensor data to fetch from the (5) IoT datalake component of the system. Once these entities are extracted, using templated query codes, specific to the IoT datalake, the system fetches the data and then maps it according to the ODA ontology to create the virtual knowledge graph and then this graph is stored inside the (4) Graph database. Inside the graph database, their is already a Base-KG that contains the system’s topological and spatial metadata and combining these with the incoming virtual knowledge graph, the system now has the required data to get the answer to the initial natural language user question. Finally, the generated SPARQL query is executed on the SPARQL endpoint of the (4) Graph Database, and the corresponding answer is then returned back to the user on the (1) Frontend, as a table, or in case of a large table with more than 50 rows, as a CSV downloadable link.

This approach maintained the semantic advantages of knowledge graphs while reducing storage overhead to manageable levels. The maximum observed virtual knowledge graph size across evaluated queries was just 179 megabytes—trivial compared to multi-terabyte full materializations. The team further optimized the generation pipeline through several technical refinements: adopting the Polars library for data processing (achieving 92.75% faster data fetching than the initial Pandas implementation), selecting N-Triples serialization format for its speed advantages, and implementing batch triple generation rather than incremental graph construction.

Natural Language Meets Graph Intelligence

The culmination of this research trajectory is EXASAGE—the first operational data analysis assistant for HPC systems [4]. The figure 4 shows the architecture diagram of EXASAGE framework. EXASAGE combines the structured power of knowledge graphs with large language models to create a natural language interface for telemetry data, enabling operators to ask questions in ordinary language and receive accurate, contextually appropriate responses.

The system architecture reflects careful integration of complementary technologies. An input validator extracts key entities from natural language questions—identifying nodes, racks, jobs, metrics, and time ranges through rule-based algorithms aligned with the domain ontology. The LLM-query generator translates these validated questions into SPARQL queries, grounded by the ontology schema and few-shot examples that demonstrate correct query patterns. Simultaneously, the virtual knowledge graph generator constructs a query-specific graph containing precisely the data needed to answer the question. A query refinement stage then corrects common LLM-generated syntax errors through pattern matching before execution.

Figure 4: EXASAGE: The first data center operational data analysis assistant, shown as a block diagram.

Evaluation across one thousand queries demonstrated remarkable effectiveness. The system achieved 93.6% accuracy in generating correct SPARQL queries and retrieving accurate answers—compared to just 25% accuracy when large language models attempted to generate equivalent NoSQL queries directly. This dramatic difference stems from the knowledge graph’s explicit encoding of relationships between data sources: information that NoSQL approaches require users to establish manually through multiple coordinated queries.

The knowledge graph approach also proved more efficient in practice. SPARQL queries generated by the system were significantly more concise, with 56.89% fewer output tokens on average, and executed faster despite the additional virtual knowledge graph construction step. Optimization efforts reduced complete end-to-end latency from 20.36 seconds to just 3.03 seconds—an 85% improvement that makes the system practical for interactive analytical sessions rather than batch processing alone [3].

Toward Universal Interoperability

The most recent advancement extends this work toward cross-system analytics through a unified ontology designed to support multiple heterogeneous HPC environments [5]. The figure 5 shows the developed ontology, comprising 164 axioms: 104 logical axioms defining semantics (e.g., domains, ranges, characteristics, inverses) and 60 declarations introducing named entities. The ontology defines 12 classes representing core HPC concepts—jobs, compute nodes, racks, sensors, users—and includes 23 object properties for inter-class relations (e.g., job-tonode, rack-to-node) and 25 data properties linking individuals to literals (e.g., timestamps, metrics). Designed for interoperability across heterogeneous HPC systems, the schema supports multiple facilities via DataCenter and HPCSystem, captures workloads through User, Job, and JobMetric, and models infrastructure layout with Rack, ComputeNode, and Position. Monitoring data is represented via Sensor and SensorReading, while temporal dynamics are abstracted with the Time class, enabling representation of time-dependent events and relationships. The Plugin class models software components involved in monitoring and analysis. The research team validated this generalized schema on the two largest publicly available operational datasets from top-ranked supercomputers: the Marconi100 dataset [1] from CINECA in Italy and F-DATA [6] from the Fugaku supercomputer at RIKEN in Japan.

Figure 5: Unified ODA Ontology

This unification required addressing fundamental differences in how facilities conceptualize and record operational data. While the Marconi100 telemetry follows a sensor-centric model where time-series readings attach to physical monitoring devices, Fugaku’s F-DATA dataset is fundamentally job-centric—recording performance metrics per job rather than per sensor. The unified ontology accommodates both paradigms through new classes that capture data center and HPC system hierarchies, user-specific workload patterns, and job-level metrics alongside traditional sensor readings.

The team also addressed storage efficiency through ontology design optimizations. By eliminating redundant classes, centralizing timestamp representations through a dedicated Time class with Unix encoding, and relocating unit specifications from individual readings to parent sensor definitions, the unified ontology achieves 38.84% storage reduction compared to the previous approach. Additional deployment configurations using blank nodes for sensor readings provide a further 26.82% reduction when global addressability of individual readings is not required.

The unified ontology was validated against 36 competency questions spanning system topology, sensor monitoring, job execution analysis, user activity patterns, scheduling efficiency, and cross-system comparative analytics. Questions that would be essentially impossible with traditional approaches—such as comparing average job execution times across different HPC systems or identifying which facility achieves better energy efficiency per job—become answerable through standardized SPARQL queries operating over the unified schema.

Implications for Operational Intelligence

This research progression—from complex NoSQL queries through unified ontologies and virtual knowledge graphs to intelligent natural language assistants—represents a fundamental shift in how data center operators can interact with telemetry data. The implications extend beyond high-performance computing to any IoT environment generating heterogeneous time-series data at scale.

The core insight is that knowledge graphs, when combined with modern language model capabilities and careful architectural optimization, can bridge the gap between massive operational datasets and actionable insights. By encoding domain semantics explicitly and leveraging the reasoning capabilities of language models, these systems enable facility managers, system administrators, and engineers to query telemetry data using questions they would naturally ask—rather than requiring mastery of database internals and query languages.

The accuracy improvements are particularly striking. A 92-93% query accuracy rate versus 25% for direct LLM-to-NoSQL approaches reflects more than incremental optimization; it demonstrates that providing structured semantic context transforms language models from unreliable query generators into effective analytical partners. The knowledge graph supplies precisely the relational information that large language models struggle to infer from unstructured database schemas.

As data centers grow in complexity to meet AI-driven computational demands, these graph-powered approaches offer a scalable path toward intelligent, accessible operational analytics. The research demonstrates that the barrier between operational data and operational insight need not remain as high as current practices suggest—and that thoughtful integration of semantic technologies with modern AI can democratize access to the intelligence hidden within facility telemetry.

Conclusion

In conclusion, thanks to the Graph-Massivizer project, CINECA and UNIBO have demonstrated that graph-based data representation is a key enabler for future autonomous and sustainable data centres: democratizing data and trasforming them from a storage overhead into actionalble insight. This research will be a valuable asset for the future development of data centres, enabling simpler, more effective, and more intelligent operational optimization.

References

[1] Borghesi, Andrea, et al. “M100 Exadata: a data collection campaign on the CINECA’s Marconi100 Tier-0 supercomputer.” Scientific Data 10.1 (2023): 288.

[2] Junaid Ahmed Khan, Martin Molan, Matteo Angelinelli, and Andrea Bartolini. 2024. ExaQuery: Proving Data Structure to Unstructured Telemetry Data in Large-Scale HPC. In Companion of the 15th ACM/SPEC International Conference on Performance Engineering (ICPE ’24 Companion). Association for Computing Machinery, New York, NY, USA, 127–134. https://doi.org/10.1145/3629527.3652898

[3] Junaid Ahmed Khan, Hiari Pizzini Cavagna, Andrea Proia, Andrea Bartolini. “From Data Center IoT Telemetry to Data Analytics Chatbots — Virtual Knowledge Graph is All You Need.” arXiv preprint, https://doi.org/10.48550/arXiv.2506.22267

[4] Khan, Junaid Ahmed, Martin Molan, and Andrea Bartolini. “EXASAGE: The First Data Center Operational Data Analysis Assistant.” Future Generation Computer Systems (2025): 108185.

[5] Khan, Junaid Ahmed, and Andrea Bartolini. “A Unified Ontology for Scalable Knowledge Graph-Driven Operational Data Analytics in High-Performance Computing Systems.” arXiv preprint arXiv:2507.06107 (2025)

[6] Antici, F., Bartolini, A., Domke, J. et al. F-DATA: A Fugaku Workload Dataset for Job-centric Predictive Modelling in HPC Systems. Sci Data 12, 1321 (2025). https://doi.org/10.1038/s41597-025-05633-1

]]>
AI and Sustainability: Bridging Innovation and Environmental Responsibility https://graph-massivizer.eu/from-big-data-to-green-data-reducing-the-environmental-impact-of-data-science-with-graph-massivizer-2/ Fri, 28 Nov 2025 11:03:02 +0000 https://graph-massivizer.wp.itec.aau.at/?p=1372 AI and Sustainability: Bridging Innovation and Environmental Responsibility

Artificial intelligence is reshaping modern society, but this technological revolution comes at an environmental cost that we can no longer ignore. Although AI offers innovative ways to tackle climate change, it also contributes to carbon emissions and the consumption of resources.

The environmental impact of AI is staggering. Research published in Nature Sustainability reveals that implementing AI servers in the United States could generate between 24 and 44 million metric tons of CO₂ equivalent emissions annually by 2030 — comparable to adding 5 to 10 million cars to American roads. The United Nations emphasises that data centres hosting AI servers produce electronic waste, consume vast amounts of water, and use huge quantities of electricity, thereby fuelling greenhouse gas emissions. Golestan Radwan, UNEP Chief Digital Officer, states that we must ensure the net effect of AI on the planet is positive before it is implemented on a large scale. AI is what researchers call “double-edged technology”.

Supercomputers and AI models have their own carbon footprint, which varies depending on the type of AI and the training methods used. However, AI can also play a key role in reducing emissions through climate change modelling and smart grid design. Microsoft Research has demonstrated that AI-based systems can integrate renewable energy more effectively into stable electrical grids and reduce carbon capture costs by accelerating the discovery of new materials.

In this critical context, Graph Massivizer shows how innovation and sustainability can go hand in hand. The project aims to improve the efficiency of data analysis and reduce the energy impact of extract, transform, and load operations. It seeks to enhance data centre energy efficiency by a factor of two and reduce greenhouse gas emissions associated with graph-organised database operations.

The Data Centre Digital Twin use case, involving CINECA and the University of Bologna, exemplifies this vision by creating a digital graph representation of CINECA’s supercomputers. This representation allows system operations to be studied and understood and enables efficiency and sustainability to be optimised for the next generation of exascale supercomputers.

Recent research from the University of Bologna and CINECA showcases AI designed to support sustainable development. The ExaQuery project proposes an innovative ontology for operational data in HPC systems that organises and queries telemetry data more efficiently, reducing computational load. This knowledge graph-based approach facilitates the identification of complex relationships between hardware components, computational jobs and performance metrics, paving the way for intelligent resource optimisation.

Building on this foundation, researchers have developed a Virtual Knowledge Graph system that provides natural language access to heterogeneous IoT data in data centres. Combining Large Language Models with Knowledge Graphs achieves 92.5% query accuracy for this system, compared to 25% for traditional LLM-to-NoSQL approaches, while reducing latency by 85%. This breakthrough demonstrates that intelligent data organisation can dramatically improve accessibility and efficiency when managing complex telemetry systems.

As Graph Massivizer highlights, although data analysis and processing have a significant environmental impact, they can also be invaluable tools for achieving environmental sustainability. The key lies in using technology responsibly by focusing on efficiency, renewable energy and intelligent computational architectures. The message is clear: AI can and must be part of the solution to the climate crisis, not the problem. Projects like Graph Massivizer show that technology’s future can be powerful and sustainable if we make the right choices today.

]]>