Due to the special characteristics of stream processing engines, the MapReduce framework and its open-source implementation Hadoop, which is mainly designed for handling batch processing tasks, are not adequate for supporting real-time stream processing. Therefore, many development efforts have created quite a few stream processing engines with the proper components to support the special requirements of the stream processing engines with parallel and scalable computation architecture. Many of them have been presented in Sakr and Gaber.
Study the reading assignments first, and write a comparison study report focusing on the following aspects:
Identify 2 different streaming processing systems, and specify the main technical features of each system.
Compare them based on the following 3 criteria:
System architecture
Performance optimization capability
Scalability
Identify the performance bottlenecks of processing the large-scale streaming dataset for each system.
Analyze the root causes for the performance bottlenecks for each system.
Propose your high-level strategies to remove these performance bottlenecks.
You need to present the justification or rationale on your point of view and should use examples to illustrate your point of view. Find at least 2 references from the Library and the Internet based on your research interests. You need at least 3 references for this report, including the 1 reference listed in this assignment.
Reference
Sakr, S., & Gaber, M. (Eds.). (2014). Large scale and big data: Processing and management. CRC Press.
Last Completed Projects
| topic title | academic level | Writer | delivered |
|---|
