Apache Hudi + Flink��ҵ��ָ��-��ƿ��

Apache Hudi + Flink��ҵ��ָ��

2024-03-12 180

��Ȩ

��Ȩ��

��ɰ��ʵ��ע��û��Է��ף��Ȩ��ԭ��У��ƿ��ӵ��Ȩ��಻�е��Ӧ��Ρ��鿴�� ƿ��û��Э�� ƿ��֪ʶ��Ȩ��ָ��ֱ��ӳ�Ϯ��ݣ��д ��ȨͶ�߱��оٱ��һ��ʵ��ɾ��Ȩ��ݡ�

��漰�Ĳ�Ʒ

ʵʱ�� Flink �棬5000CU*H 3��

��飺 Apache Hudi + Flink��ҵ��ָ��

��Apache Hudi��ϲ��Flink��Ļ��ʵ�֣�HUDI-1327��ζ�� Hudi ��ʼ֧�� Flink ��档�кܶ�С��ڽ��Ⱥ��ѯ Hudi on Flink ��ʹ��ƣ��ﲻ��ʵ��ʾһ�ѣ��ƪ��¡�

��ǰ Flink �汾��Hudi��ֻ֧�ֶ�ȡ Kafka ��ݣ�Sink�� COW��COPY_ON_WRITE�� ͵� Hudi ��У��ܻ��ڼ��С�

��Ǽ�Ҫ��δ� Kafka ��ȡ��д��Hudi��

1. ��

��ڻ�û��ʽ��, ��Ҫ��Github��Դ��д��

git clone https://github.com/apache/hudi.git && cd hudimvn clean package -DskipTests

Windows ϵͳ�û��ʱ�ᱨ��´��

[ERROR] Failed to execute goal org.codehaus.mojo:exec-maven-plugin:1.6.0:exec (Setup HUDI_WS) on project hudi-integ-test: Command execution failed. Cannot run program "\bin\bash" (in directory "D:\github\hudi\hudi-integ-test"): CreateProcess error=2, ϵͳ�Ҳ���ָ�����ļ��� -> [Help 1][ERROR][ERROR] To see the full stack trace of the errors, re-run Maven with the -e switch.[ERROR] Re-run Maven using the -X switch to enable full debug logging.[ERROR][ERROR] For more information about the errors and possible solutions, please read the following articles:[ERROR] [Help 1] http://cwiki.apache.org/confluence/display/MAVEN/MojoExecutionException[ERROR][ERROR] After correcting the problems, you can resume the build with the command[ERROR]   mvn <goals> -rf :hudi-integ-test

�� hudi-integ-test ģ��һ��bash�ű��޷�ִ�е��µĴ��ǿ��԰��ע�͵��

�޸�D:\github\hudi\pom.xml��pom�ļ�

<modules>    <module>hudi-common</module>    <module>hudi-cli</module>    <module>hudi-client</module>    <module>hudi-hadoop-mr</module>    <module>hudi-spark</module>    <module>hudi-timeline-service</module>    <module>hudi-utilities</module>    <module>hudi-sync</module>    <module>packaging/hudi-hadoop-mr-bundle</module>    <module>packaging/hudi-hive-sync-bundle</module>    <module>packaging/hudi-spark-bundle</module>    <module>packaging/hudi-presto-bundle</module>    <module>packaging/hudi-utilities-bundle</module>    <module>packaging/hudi-timeline-server-bundle</module>    <module>docker/hoodie/hadoop</module><!--    <module>hudi-integ-test</module>--><!--    <module>packaging/hudi-integ-test-bundle</module>-->    <module>hudi-examples</module>    <module>hudi-flink</module>    <module>packaging/hudi-flink-bundle</module>  </modules>

�ٴ�ִ�� mvn clean package -DskipTests, ִ�гɹ��ҵ��jar : D:\github\hudi\packaging\hudi-flink-bundle\target\hudi-flink-bundle_2.11-0.6.1-SNAPSHOT.jar (��HudiԴ��D:\github\ ·��£��Ҹ��Լ�ʵ��·��һ��)

�� hudi-flink-bundle_2.11-0.6.1-SNAPSHOT.jar ��Ҫʹ�õ�flink�ͻ��ˣ��ԭ�� hudi-utilities-bundle_2.11-x.x.x.jar ��

2. ��ν��

�м��ش��Ĳ��£�

?--kafka-topic ��Kafka ��?--kafka-group-id ��?--kafka-bootstrap-servers : Kafka brokers?--target-base-path : Hudi ��·��?--target-table ��Hudi ��?--table-type ��Hudi ��?--props : ��

��Բο� org.apache.hudi.HoodieFlinkStreamer.Config��ÿ��н��

3. ��׼��嵥

1.Kafka ��⣬��2.jar�ϴ��3.schema �ļ�4.Hudi��ļ�

ע��Լ��ð��ļ��ŵ��ʵĵط��ߵ� hudi-conf.properties��schem.avsc�ļ��ϴ��HDFS��

-rw-r--r-- 1 user user      592 Nov 19 09:32 hudi-conf.properties-rw-r--r-- 1 user user 39086937 Nov 30 15:51 hudi-flink-bundle_2.11-0.6.1-SNAPSHOT.jar-rw-r--r-- 1 user user 1410 Nov 17 17:52 schema.avsc

hudi-conf.properties��

hoodie.datasource.write.recordkey.field=uuidhoodie.datasource.write.partitionpath.field=tsbootstrap.servers=xxx:9092hoodie.deltastreamer.keygen.timebased.timestamp.type=EPOCHMILLISECONDShoodie.deltastreamer.keygen.timebased.output.dateformat=yyyy/MM/ddhoodie.datasource.write.keygenerator.class=org.apache.hudi.keygen.TimestampBasedAvroKeyGeneratorhoodie.embed.timeline.server=falsehoodie.deltastreamer.schemaprovider.source.schema.file=hdfs://olap/hudi/test/config/flink/schema.avschoodie.deltastreamer.schemaprovider.target.schema.file=hdfs://olap/hudi/test/config/flink/schema.avsc

schema.avsc��

{  "type":"record",  "name":"stock_ticks",  "fields":[{     "name": "uuid",     "type": "string"  }, {     "name": "ts",     "type": "long"  }, {     "name": "symbol",     "type": "string"  },{     "name": "year",     "type": "int"  },{     "name": "month",     "type": "int"  },{     "name": "high",     "type": "double"  },{     "name": "low",     "type": "double"  },{     "name": "key",     "type": "string"  },{     "name": "close",     "type": "double"  }, {     "name": "open",     "type": "double"  }, {     "name": "day",     "type":"string"  }]}

4. ��

/opt/flink-1.11.2/bin/flink run -c org.apache.hudi.HoodieFlinkStreamer -m yarn-cluster -d -yjm 1024 -ytm 1024 -p 4 -ys 3 -ynm hudi_on_flink_test hudi-flink-bundle_2.11-0.6.1-SNAPSHOT.jar --kafka-topic hudi_test_flink --kafka-group-id hudi_on_flink --kafka-bootstrap-servers xxx:9092 --table-type COPY_ON_WRITE --target-base-path hdfs://olap/hudi/test/data/hudi_on_flink --target-table hudi_on_flink  --props hdfs://olap/hudi/test/config/flink/hudi-conf.properties --checkpoint-interval 3000 --flink-checkpoint-path hdfs://olap/hudi/hudi_on_flink_cp

�鿴��ҳ�棬��Ѿ��

��Hdfs·��Ѿ��һ��ձ��Hudi�Զ��

�� topic �з��ݣ�� 900 ��д�� Producer �Ͳ��ˣ�

��ǲ�һ�½��

@Test  public void query() {    spark.read().format("hudi")        .load(basePath + "/*/*/*/*")        .createOrReplaceTempView("tmp_view");    spark.sql("select * from tmp_view limit 2").show();    spark.sql("select count(1) from tmp_view").show();  }

+-------------------+--------------------+--------------------+----------------------+--------------------+--------------------+-------------+--------------------+----+-----+-------------------+------------------+------+------------------+-------------------+---+|_hoodie_commit_time|_hoodie_commit_seqno|  _hoodie_record_key|_hoodie_partition_path|   _hoodie_file_name|                uuid|           ts|              symbol|year|month|               high|               low|   key|             close|               open|day|+-------------------+--------------------+--------------------+----------------------+--------------------+--------------------+-------------+--------------------+----+-----+-------------------+------------------+------+------------------+-------------------+---+|     20201130162542| 20201130162542_0_20|01e11b9c-012a-461...|            2020/10/29|c8f3a30a-0523-4c8...|01e11b9c-012a-461...|1603947341061|12a-4614-89c3-f62...| 120|   10|0.45757580489415417|0.0816472025173598|01e11b|0.5795817262998396|0.15864898816336837|  1||     20201130162542| 20201130162542_0_21|22e96b41-344a-4be...|            2020/10/29|c8f3a30a-0523-4c8...|22e96b41-344a-4be...|1603921161580|44a-4be2-8454-832...| 120|   10| 0.6200960168557579| 0.946080636091312|22e96b|0.6138608980526853| 0.5445994550724997|  1|+-------------------+--------------------+--------------------+----------------------+--------------------+--------------------+-------------+--------------------+----+-----+-------------------+------------------+------+------------------+-------------------+---+

+--------+|count(1)|+--------+|     900|+--------+

5. �ܽ�

��ļ�Ҫ��ʹ�� Flink ��潫��д��Hudi��Ĺ��̡��Ҫ��ִ��jar��ܡ�Schema��á�Hudi��õȲ��

��ʵ��ѧϰ

��Hologres��תһվʽʵʱ�ֿ�

��ð��MaxCompute��ʵʱ��Flink�ͽ��ʽ��Hologres��ߡ�ʵʱ��ںϷ��ݴ��Ӧ�á�

Linux��ŵ��ͨ

��׿γ��Ǵ��ſ�ʼ��Linuxѧϰ�γ̣��ʺϳ�ѧ��Ķ��ǳ����ḻ��ͨ��׶��Ҫ�漰��ϵͳ��Լ��г��õĸ��ַ��Ӧ�á��Ż��ʹ��ѧԱ��ֻҪ�ܹ��ְ��½ڶ�ѧ�꣬Ҳһ��ǳ��

Apache Hudi + Flink��ҵ��ָ��

1. ��

2. ��ν��

3. ��׼��嵥

4. ��

5. �ܽ�

��

��

��ؿγ�

��ص��

��ʵ�鳡��

�Ƽ��

Apache Hudi + Flink��ҵ����ָ��

1. ���

2. ��ν���

3. ����׼���嵥

4. ��������

5. �ܽ�

��������

��������

��ؿγ�

��ص�����

���ʵ�鳡��

�Ƽ�����

Apache Hudi + Flink��ҵ��ָ��

1. ��

2. ��ν��

3. ��׼��嵥

4. ��

��

��

��ص��

��ʵ�鳡��

�Ƽ��