Wu, Y.; Tan, K. L.: ChronoStream: Elastic stateful stream
computation in the cloud. IEEE ICDE, 2015.
弾力性の実現に関する先行研究
ただし別の論文の指摘によると、同期に関する課題がある
データフロー変更には対応していない。あくまでオペレータのマイグレーションのみ。
Heinze, T.; Pappalardo, V.; Jerzak, Z.; Fetzer, C.: Auto-scaling
techniques for elastic data stream processing. In: IEEE ICDE Workshops.
2014. および、 Heinze, T.; Ji, Y.; Roediger, L.; Pappalardo, V.;
Meister, A.; Jerzak, Z.; Fetzer, C.: FUGU: Elastic Data Stream
Processing with Latency Constraints. IEEE Data Eng. Bull., 2015.
オートスケールのタイミングを判断する。オンライン機械学習を利用。
簡単なマイグレーションのシナリオを想定。
Nasir, M.; Morales, G.; Kourtellis, N.; Serafini, M.: When Two
Choices Are not Enough: Balancing at Scale in Distributed Stream
Processing. CoRR, abs/1510.05714, 2015.
並列度の調整に関する先行研究
特にホットキーが存在する場合、そこにオペレータインスタンスを割り当てるように動く。
データフロー変更などには対応しない。
Mai, L.; Zeng, K.; Potharaju, R.; Xu, L.; Venkataraman, S.; Costa,
P.; Kim, T.; Muthukrishnan, S.; Kuppa, V.; Dhulipalla, S.; Rao, S.: Chi:
A Scalable and Programmable Control Plane for Distributed Stream
Processing Systems. VLDB, 2018.
Apache Edgent is a programming model and micro-kernel style runtime
that can be embedded in gateways and small footprint edge devices
enabling local, real-time, analytics on the continuous streams of data
coming from equipment, vehicles, systems, appliances, devices and
sensors of all kinds (for example, Raspberry Pis or smart phones).
Working in conjunction with centralized analytic systems, Apache Edgent
provides efficient and timely analytics across the whole IoT ecosystem:
from the center to the edge.
public class ImpressiveEdgentExample { public static void main(String[] args) { DirectProvider provider = new DirectProvider(); Topology top = provider.newTopology();
IotDevice iotConnector = IotpDevice.quickstart(top, "edgent-intro-device-2"); // open https://quickstart.internetofthings.ibmcloud.com/#/device/edgent-intro-device-2
The connector uses and includes components from the Kafka 0.8.2.2
release. It has been successfully tested against kafka_2.11-0.10.1.0 and
kafka_2.11-0.9.0.0 server as well. For more information about Kafka see
http://kafka.apache.org
Rettig, L., Khayati, M., Cudré-Mauroux, P., Piórkowski, M., 2015. Online anomaly detection over big data streams. In: IEEE International Conference on Big Data (Big Data 2015), IEEE, Santa Clara, USA, pp. 11131122
Zhao, X., Garg, S., Queiroz, C., Buyya, R., 2017. Software Architecture for Big Data and the Cloud, Elsevier Morgan Kaufmann. Ch. A Taxonomy and Survey of Stream Processing Systems.
本論文は以下の構成。
2章
ビッグデータエコシステム。オンラインデータ処理のアーキテクチャ。
3章
既存のエンジン。ストリームデータ処理の他のソリューション
4章
マネージドクラウドソリューション
5章
既存の技術がどのようにリソースの弾力性を実現しようとしているのか
6章
複数のインフラストラクチャを組み合わせる
7章
将来の話
2. Background and architecture
2.1. Online data processing
architecture
ここでは「オンライン」という単語を用いているが、以下のように定義して用いている。
1 2
Similar to Boykin et al., hereafter use the term online to mean that “data are processed as they are being generated”.
Sajjad, H.P., Danniswara, K., Al-Shishtawy, A., Vlassov, V., 2016. SpanEdge: Towards unifying stream processing over central and near-the-edge data centers. In: IEEE/ ACM Symposium on Edge Computing (SEC), pp. 168178.
SPEとして。
1 2 3
Chan, S., 2016. Apache quarks, watson, and streaming analytics: Saving the world, one smart sprinkler at a time. Bluemix Blog (June).URL 〈https://www.ibm.com/blogs/ bluemix/2016/06/better-analytics-with-apache-quarks/〉
1 2 3 4
Pisani, F., Brunetta, J.R., do Rosario, V.M., Borin, E., 2017. Beyond the fog: Bringing cross-platform code execution to constrained iot devices. In: Proceedings of the 29th International Symposium on Computer Architecture and High Performance Computing (SBAC-PAD 2017), Campinas, Brazil, pp. 1724.
クラウドとエッジ。
1 2 3
Hirzel, M., Schneider, S., Gedik, B., An, S.P.L., 2017. extensible language for distributed stream processing. ACM Trans. Program. Lang. Syst. 39 (1), 5:15:39. http:// dx.doi.org/10.1145/3039207, (URL 〈http://doi.acm.org/10.1145/3039207〉).
Satzger, B., Hummer, W., Leitner, P., Dustdar, S., 2011. Esc: Towards an elastic stream computing platform for the cloud. In: IEEE International Conference on Cloud Computing (CLOUD), pp. 348–355.
ただし最近は Ning
Wang
のように見える。彼はもともと2013年あたりまでGoogleでYouTubeに携わっていたようだ。
READMEによる「Update」
2019/7/8時点のREADMEによると、Mesos in AWS、Mesos/Aurora in
AWS、ローカル(ラップトップ)の上でネイティブ動作するようになった。
またApache REEFを用いてApache YARN上で動作するように試みている。 slurm
にも対応しようとしているとのこと。
this.accumulator = new RecordAccumulator(logContext, config.getInt(ProducerConfig.BATCH_SIZE_CONFIG), this.compressionType, lingerMs(config), retryBackoffMs, deliveryTimeoutMs, metrics, PRODUCER_METRIC_GROUP_NAME, time, apiVersions, transactionManager, new BufferPool(this.totalMemorySize, config.getInt(ProducerConfig.BATCH_SIZE_CONFIG), metrics, time, PRODUCER_METRIC_GROUP_NAME));