Person: Acacio Sánchez, Manuel Eugenio
Loading...
Name
Acacio Sánchez, Manuel Eugenio
publication.page.department
Universidad de Murcia. Departamento de Ingeniería y Tecnología de Computadores
- Publications
- item.page.relationships.isSecondaryAuthorOfPublication
- item.page.relationships.isDirectorOfPublication
Search Results
Now showing 1 - 2 of 2
- PublicationRestrictedQuCo: efficient and flexible hardware-driven automatic configuration of tile transfers in GPUs(IEEE Computer Society Press, 2025-12-16) Meseguer, Nicolás; Xu, Daoxuan; Sun, Yifan; Pellauer, Michael; Abellán Miguel, José Luis; Acacio Sánchez, Manuel Eugenio; Ingeniería y Tecnología de Computadores; Facultades de la UMU::Facultad de InformáticaThe growing complexity and parallelism demands of modern GPU workloads have driven architectural innovations toward \emph{asynchronous tile transfers} (ATTs) to overlap computation and data movement. While ATT units such as the NVIDIA’s Tensor Memory Accelerator (TMA) introduce high-throughput memory transfers, programmers must deal with wavefront specialization, select tile sizes, queue slots, and synchronization primitives, all of which are hardware-specific and workload-dependent. Existing GPU libraries fall short—offering limited ATT support and configurability—so developers still resort to manual exploration of this vast parameter space, which is laborious, error-prone, and fundamentally limits performance portability across GPUs. In this work, we present QuCo (Queue Configurator), a single lightweight hardware unit embedded in the GPU that fully automates the ATT configuration process. Inspired by Blackwell GPU design, QuCo includes a compact \mbox{RISC-V} processor, small memory structures for instructions and data, and a GPU Specification Table (GST) storing key architectural parameters. Using the GST and workload characteristics, along with built-in heuristics, QuCo computes optimal queue configurations at kernel launch. This relieves the programmer of the tedious, time-consuming task of tuning and offline profiling, while simultaneously increasing post-compilation performance portability.
- PublicationOpen AccessFlexagon: a multi-dataflow sparse-sparse matrix multiplication accelerator for efficient DNN processing(Association for Computing Machinery, 2023-03-25) Garg, Raveesh; Pellauer, Michael; Krishna, Tushar; Muñoz Martínez, Francisco; Abellán Miguel, José Luis; Acacio Sánchez, Manuel Eugenio; Ingeniería y Tecnología de Computadores; Facultad de InformáticaSparsity is a growing trend in modern DNN models.Existing Sparse-Sparse Matrix Multiplication (SpMSpM) accel-erators are tailored to a particular SpMSpM dataflow (i.e., InnerProduct, Outer Product or Gustavson’s), which determines theiroverall efficiency. We demonstrate that this static decision inher-ently results in a suboptimal dynamic solution. This is becausedifferent SpMSpM kernels show varying features (i.e., dimensions,sparsity pattern, sparsity degree), which makes each dataflow bettersuited to different data sets.In this work we present Flexagon, the first SpMSpM reconfig-urable accelerator that is capable of performing SpMSpM computa-tion by using the particular dataflow that best matches each case.Flexagon accelerator is based on a novel Merger-Reduction Net-work (MRN) that unifies the concept of reducing and merging inthe same substrate, increasing efficiency. Additionally, Flexagonalso includes a new L1 on-chip memory organization, specificallytailored to the different access characteristics of the input and out-put compressed matrices. Using detailed cycle-level simulation ofcontemporary DNN models from a variety of application domains,we show that Flexagon achieves average performance benefits of4.59×, 1.71×, and 1.35×with respect to the state-of-the-art SIGMA-like, SpArch-like and GAMMA-like accelerators (265%, 67%, and18%, respectively, in terms of average performance/area efficiency).
Ir a Estadísticas
Sin licencia Creative Commons.






