我目前有flink的设置,有一个在emr上运行的工作,我现在正试图通过向prometheus发送度量来添加监控。
我遇到了一个问题与运行电子病历Flink。我正在使用terraform来提供emr(我在下载和运行作业之后运行ansible)。开箱即用,看起来emr的flink发行版并不包括可选的jar(flink metrics prometheus、flink cep等)。
看看Flink的文件,上面说
“为了使用这个记者,你必须复制 /opt/flink-metrics-prometheus-1.6.1.jar
进入 /lib
flink发行版的文件夹“https://ci.apache.org/projects/flink/flink-docs-release-1.6/monitoring/metrics.html#prometheuspushgateway-orgapacheflinkmetricsprometheusprometheuspushgatewayreporter公司
但是当登录到emr主节点时,/etc/flink或/usr/lib/flink都没有名为 opts
我看不见 flink-metrics-prometheus-1.6.1.jar
任何地方。
我知道flink还有其他可选的lib,如果你想使用它们,比如flink cep,你通常必须复制它们,但是我不知道在使用emr时如何做到这一点。
这是我得到的例外,我认为是因为它在类路径中找不到metrics jar。
java.lang.ClassNotFoundException: org.apache.flink.metrics.prometheus.PrometheusPushGatewayReporter
at java.net.URLClassLoader.findClass(URLClassLoader.java:382)
at java.lang.ClassLoader.loadClass(ClassLoader.java:424)
at sun.misc.Launcher$AppClassLoader.loadClass(Launcher.java:349)
at java.lang.ClassLoader.loadClass(ClassLoader.java:357)
at java.lang.Class.forName0(Native Method)
at java.lang.Class.forName(Class.java:264)
at org.apache.flink.runtime.metrics.MetricRegistryImpl.<init>(MetricRegistryImpl.java:144)
at org.apache.flink.runtime.entrypoint.ClusterEntrypoint.createMetricRegistry(ClusterEntrypoint.java:419)
at org.apache.flink.runtime.entrypoint.ClusterEntrypoint.initializeServices(ClusterEntrypoint.java:276)
at org.apache.flink.runtime.entrypoint.ClusterEntrypoint.runCluster(ClusterEntrypoint.java:227)
at org.apache.flink.runtime.entrypoint.ClusterEntrypoint.lambda$startCluster$0(ClusterEntrypoint.java:191)
at java.security.AccessController.doPrivileged(Native Method)
at javax.security.auth.Subject.doAs(Subject.java:422)
at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1844)
at org.apache.flink.runtime.security.HadoopSecurityContext.runSecured(HadoopSecurityContext.java:41)
at org.apache.flink.runtime.entrypoint.ClusterEntrypoint.startCluster(ClusterEntrypoint.java:190)
at org.apache.flink.yarn.entrypoint.YarnSessionClusterEntrypoint.main(YarnSessionClusterEntrypoint.java:137)
地形中的emr资源
resource "aws_emr_cluster" "emr_flink" {
name = "ce-emr-flink-arn"
release_label = "emr-5.20.0" # 5.21.0 is not found, could be a region thing
applications = ["Flink"]
ec2_attributes {
key_name = "ce_test"
subnet_id = "${aws_subnet.ce_test_subnet_public.id}"
instance_profile = "${aws_iam_instance_profile.emr_profile.arn}"
emr_managed_master_security_group = "${aws_security_group.allow_all_vpc.id}"
emr_managed_slave_security_group = "${aws_security_group.allow_all_vpc.id}"
additional_master_security_groups = "${aws_security_group.external_connectivity.id}"
additional_slave_security_groups = "${aws_security_group.external_connectivity.id}"
}
ebs_root_volume_size = 100
master_instance_type = "m4.xlarge"
core_instance_type = "m4.xlarge"
core_instance_count = 2
service_role = "${aws_iam_role.iam_emr_service_role.arn}"
configurations_json = <<EOF
[
{
"Classification": "flink-conf",
"Properties": {
"parallelism.default": "8",
"state.backend": "RocksDB",
"state.backend.async": "true",
"state.backend.incremental": "true",
"state.savepoints.dir": "file:///savepoints",
"state.checkpoints.dir": "file:///checkpoints",
"web.submit.enable": "true",
"metrics.reporter.promgateway.class": "org.apache.flink.metrics.prometheus.PrometheusPushGatewayReporter",
"metrics.reporter.promgateway.host": "${aws_instance.monitoring.private_ip}",
"metrics.reporter.promgateway.port": "9091",
"metrics.reporter.promgateway.jobName": "ce-test",
"metrics.reporter.promgateway.randomJobNameSuffix": "true",
"metrics.reporter.promgateway.deleteOnShutdown": "false"
}
}
]
EOF
}
我怀疑我可能需要在引导阶段下载jar,但是我想先检查一下,看看是否有这样的例子
2条答案
按热度按时间mi7gmzs61#
我选择了emr发行版emr-5.24.0,并使用influxdb.jar监视suceed。
我已将.jar文件复制到
/usr/lib/flink/lib
文件夹并使用以下bash命令(具有sudo权限)重新启动flink集群。我想你可以用普罗米修斯的相同步骤来解决你的问题
bsxbgnwa2#
我没有使用terraform,但是请注意,您通常需要在emr中对主服务器和从服务器进行配置(设置jar)。找出emr认为jar应该去哪里的一种方法是在作业运行时登录到从属服务器,do
ps auxwww | grep java
,找到TaskManager
在这个过程中,查看类路径启动时添加到类路径中的jar,并找到它们在服务器上的位置。或者至少在过去对我有用。