CBL - Campus del Baix Llobregat

Projecte llegit

Títol: Distributed Computing Infrastructure for AI Workloads


Estudiants que han llegit aquest projecte:


Director/a: SPADARO, SALVATORE

Departament: TSC

Títol: Distributed Computing Infrastructure for AI Workloads

Data inici oferta: 16-10-2025     Data finalització oferta: 16-05-2026



Estudis d'assignació del projecte:
    MU AI4CI
    MU EM CODAS 1
    MU EM CODAS 2
    MU MASTEAM 2015
Tipus: Individual
 
Lloc de realització: EETAC
 
Paraules clau:
Optical networks, cloud services, deterministic optical systems, Distributed Computing Infrastructure
 
Descripció del contingut i pla d'activitats:
The development of emerging applications and use cases such as, among others, artificial intelligence (AI), autonomous and connected vehicles, virtual reality and industrial factories digitalization, has a significant impact on the requirements posed not only to transmission networks but also in terms of computing power. Current efforts towards design and implementation of 6G systems are envisioned to enable future networks to provide support to such new ecosystem of applications, services and use cases. To meet these requirements, the concept of Telecom Cloud Continuum, that relies on the integration of cloud computing and telecommunications networks to provide telecom services, arise.
The Master thesis objectives will be to define suitable architectural solutions to enable AI-driven Telecom Cloud ecosystems.
 
Overview (resum en anglès):
Artificial intelligence workloads are changing the way computing infrastructure
is designed. Large-scale training and low-latency inference require increasing
computing capacity, high bandwidth and predictable communication
performance. While Scale-Up and Scale-Out can extend the resources
available inside a single data center, they are eventually limited by factors such
as power, cooling, physical space and network scalability. This motivates the
use of geographically distributed Edge and Cloud data centers connected
through a Telecom Cloud Network.
This thesis studies how computing and network resources can be coordinated
to support AI workloads across multiple sites. It reviews the evolution towards
distributed AI infrastructure, identifies the network requirements imposed by
these workloads and analyses the limitations of single-site orchestration.
Based on this analysis, an end-to-end architecture is proposed in which Edge
and Cloud data centers expose abstract computing-resource information, while
the Telecom Cloud Network exposes the connectivity available between sites.
An E2E Infrastructure Orchestrator combines both views to select a feasible
compute-and-network combination. The network integrates packet connectivity
with programmable optical transport, using shared packet paths when sufficient
and reconfigurable optical lightpaths when higher sustained capacity or
reserved connectivity is required.
Two illustrative cases, latency-sensitive inference and distributed training, are
used to show how workload characteristics affect placement and connectivity
decisions. For example, transferring 1 GB of synchronization data within 100
ms requires an average rate of 80 Gb/s. The analysis indicates that, in
distributed AI infrastructures, the network cannot be treated as an unlimited
and always-available resource; computing-site selection and network-resource
selection should therefore be coordinated as part of the same orchestration
process.


© CBLTIC Campus del Baix Llobregat - UPC