GB/T 44216-2024 in English
VALIDInformation technology—Big data—Techinical requirements for integrated batch and streaming computing
- Issued on:2024-07-24
- Implemented on:2025-02-01
- File Format:PDF
- Delivery:Via email within 5 business days
$301.00
《GB/T 44216-2024信息技术 大数据 批流融合计算技术要求》由TC28(全国信息技术标准化技术委员会)归口,主管部门为国家标准委。
Introduction
National Standard of the People's Republic of China GB/T 44216-2024
Technical Requirements for Batch-Stream Fusion Computing of Big Data in Information Technology
Foreword
This standard is drafted in accordance with the provisions of GB/T 1.1-2020 "Guidelines for Standardization Work Part 1: Structure and Drafting Rules for Standardization Documents". Please note that some contents of this standard may involve patents. The issuing agency of this standard does not assume the responsibility for identifying patents.
This standard is proposed and managed by the National Technical Committee for Information Technology Standardization (SAC/TC 28). The main drafting units include Alibaba Cloud Computing Co., Ltd., China Electronics Technology Standardization Institute and many other companies and research institutions, jointly promoting the standardization development of batch-stream fusion computing technology.
Introduction
With the growth of data volume, distributed computing mode has become the mainstream architecture for big data processing and computing. Batch computing technology and stream computing technology often need to work together in practical applications due to their different programming models and interfaces. However, simple superposition will lead to increased development and maintenance costs.
Therefore, unified batch-stream fusion computing technology has become an important development trend in the field of big data, which can support computing frameworks such as batch processing and stream processing at the same time, and achieve efficient resource utilization and seamless connection of tasks.
| Technical requirements dimension | Batch processing characteristics | Stream processing characteristics | Batch-stream fusion advantages |
|---|---|---|---|
| Resource management | Supports large-scale parallel computing, suitable for batch data processing. | Strong real-time performance, suitable for high-speed data stream processing. | Uniform scheduling, dynamic allocation of resources, and improved utilization. |
| Computing framework | Based on the MapReduce model, suitable for processing complete record sets. | Event-driven, supports complex event processing (CEP) and real-time analysis. | Unified computing framework, supports batch and stream mixed tasks. |
| SQL interface | Supports standard SQL aggregation, connection and sorting operations. | Supports window aggregation and time processing functions. | Unified interface, compatible with batch and stream data types and operations. |
Technical Requirements Interpretation
6.1 Unified Resource Management
Batch-stream fusion computing technology must have unified resource scheduling and allocation capabilities, and support dynamic adjustment of heterogeneous resources (CPU, memory, GPU). Through the combination of static and dynamic strategies, task priority scheduling and elastic resource preemption are realized to ensure resource isolation and sharing in a multi-tenant environment.
6.2 Unified Computing Framework
The unified computing framework must support advanced functions such as complex event processing (CEP), machine learning, and graph computing. Through window segmentation, watermark operation, and trigger mechanism, efficient processing of streaming data is achieved. At the same time, it supports DAG job process orchestration and multiple time definitions to meet the needs of different scenarios.
6.3 Unified SQL Interface
The unified SQL interface needs to cover batch and stream data types (such as variable-length strings, integers, timestamps, etc.), and support window aggregation and multiple connection operations. The flexibility of data processing is improved through the extension of user-defined functions (UDF, UDTF, UDAF) and system functions.
6.4 Unified API
The unified API describes the abstract interface of batch and stream jobs and supports the docking of multiple data storage systems. Through connectors and user-defined operators, compatibility with third-party storage systems and computing engines is achieved.
Implementation Suggestions
Resource Planning and Allocation
In actual deployment, computing resources need to be reasonably planned according to business needs and data scale. It is recommended to adopt a multi-tenant mode to ensure elastic expansion and isolation of resources.
System Architecture Design
Design the system architecture based on a unified batch-stream fusion computing framework, and give priority to engines that support complex event processing and machine learning capabilities. Optimize task scheduling through queue resource management and dynamic resource preemption mechanism.
Performance Monitoring and Tuning
Regularly monitor job status, data skew, and resource utilization, and adjust computing parameters in a timely manner. It is recommended to use a unified metric interface to implement custom monitoring of batch-stream tasks.

Loading PDF document...
Error loading PDF. Please make sure the file is valid and try again.
We also recommend
-

GB/T 44272-2024 in English
Information technology—Open source—Open source license framework
2024-08-23 -

GB/Z 191-2026 in English
Intelligent computing—Technical framework of privacy-preserving computation
2026-07-02 -

GB/T 46883-2025 in English
Industrial internet platform—Specification for digital service of industrial park
2025-12-31 -

GB/Z 20539-2006 in English
Modeling guide for business process and information in e-Business
2006-09-18 -

GB/T 41575-2022 in English
Classification and code of internet unhealthy content for minors
2022-07-11 -

GB/T 37721-2019 in English
Information technology—Functional requirements for big data analytic systems
2019-08-30