arrow
Return

Building and Operating a Large-Scale Enterprise Data Analytics Platform

delete2021-02-01
delete6
delete
OA
AI
D
Daniel Bauer
F
Florian Froese
L
Luis Garcés-Erice
C
Chris Giblin
A
Abdel Labbi
Z
Zoltán András Nagy
N
Niels Pardon
S
Seán Rooney *
P
Peter Urbanetz
P
Pascal Vetsch
A
Andreas Wespi
DOI:10.1016/j.bdr.2020.100181delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Over the last three years we have been running a large-scale data processing platform for applying analytics to corporate data at scale on an OpenStack private cloud instance. Our platform makes a wide variety of corporate data assets, such as sales, marketing, customer information, as well as data from less conventional sources such as weather, news and social media available for analytics purposes to hundreds of globally distributed teams across the company. We control every layer in the stack from the processing engines down to the hardware. Here we report our experiences in building and operating such a system. We describe our technical choices and describe how they evolved as we observed the actual workloads created by users. (C) 2020 The Authors. Published by Elsevier Inc.
Keywords:
Hybrid cloud
Datalake
Storage
Ingestion
SQL/Hadoop
Governance
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Big Data Research cover
Big Data Research
IF:
4.2
Papers:
406
Citations:
1.1K

Organization

No organization information available