千家信息网

CDH: unable to create new native thread

发表于:2025-01-26 作者:千家信息网编辑
千家信息网最后更新 2025年01月26日,发现问题CDH-4.7.1 NameNode is down启动NameNode报错如下,无法创建新的线程,可能是使用的线程数超过max user processes设定的阈值2018-08-26 0
千家信息网最后更新 2025年01月26日CDH: unable to create new native thread

发现问题

CDH-4.7.1 NameNode is down

启动NameNode报错如下,无法创建新的线程,可能是使用的线程数超过max user processes设定的阈值

2018-08-26 08:44:00,532 INFO org.apache.hadoop.http.HttpServer: Jetty bound to port 500702018-08-26 08:44:00,532 INFO org.mortbay.log: jetty-6.1.26.cloudera.42018-08-26 08:44:00,773 WARN org.apache.hadoop.security.authentication.server.AuthenticationFilter: 'signature.secret' configuration not set, using a random value as secret2018-08-26 08:44:00,812 INFO org.mortbay.log: Started SelectChannelConnector@alish2-dataservice-01.mypna.cn:500702018-08-26 08:44:00,813 INFO org.apache.hadoop.hdfs.server.namenode.NameNode: Web-server up at: alish2-dataservice-01.mypna.cn:500702018-08-26 08:44:00,814 INFO org.apache.hadoop.ipc.Server: IPC Server Responder: starting2018-08-26 08:44:00,815 INFO org.apache.hadoop.ipc.Server: IPC Server listener on 8020: starting2018-08-26 08:44:00,828 INFO org.apache.hadoop.ipc.Server: IPC Server Responder: starting2018-08-26 08:44:00,828 INFO org.apache.hadoop.ipc.Server: IPC Server listener on 8022: starting2018-08-26 08:44:00,839 FATAL org.apache.hadoop.hdfs.server.namenode.NameNode: Exception in namenode joinjava.lang.OutOfMemoryError: unable to create new native threadat java.lang.Thread.start0(Native Method)at java.lang.Thread.start(Thread.java:714)at org.apache.hadoop.ipc.Server.start(Server.java:2057)at org.apache.hadoop.hdfs.server.namenode.NameNodeRpcServer.start(NameNodeRpcServer.java:303)at org.apache.hadoop.hdfs.server.namenode.NameNode.startCommonServices(NameNode.java:497)at org.apache.hadoop.hdfs.server.namenode.NameNode.initialize(NameNode.java:459)at org.apache.hadoop.hdfs.server.namenode.NameNode.(NameNode.java:621)at org.apache.hadoop.hdfs.server.namenode.NameNode.(NameNode.java:606)at org.apache.hadoop.hdfs.server.namenode.NameNode.createNameNode(NameNode.java:1177)at org.apache.hadoop.hdfs.server.namenode.NameNode.main(NameNode.java:1241)2018-08-26 08:44:00,851 INFO org.apache.hadoop.util.ExitUtil: Exiting with status 1


日志内容如下,检查DNS没有问题,这里没有太多参考意义

#cat /var/log/cloudera-scm-agent/cloudera-scm-agent.log[26/Aug/2018 07:30:23 +0000] 4589 MainThread agent        INFO     PID '19586' associated with process '1724-hdfs-NAMENODE' with payload 'processname:1724-hdfs-NAMENODE groupname:1724-hdfs-NAMENODE from_state:RUNNING expected:0 pid:19586' exited unexpectedly[26/Aug/2018 07:45:06 +0000] 4589 Monitor-HostMonitor throttling_logger ERROR    (29 skipped) Failed to collect java-based DNS namesTraceback (most recent call last):  File "/usr/lib64/cmf/agent/src/cmf/monitor/host/dns_names.py", line 53, in collect    result, stdout, stderr = self._subprocess_with_timeout(args, self._poll_timeout)  File "/usr/lib64/cmf/agent/src/cmf/monitor/host/dns_names.py", line 42, in _subprocess_with_timeout    return subprocess_with_timeout(args, timeout)  File "/usr/lib64/cmf/agent/src/cmf/monitor/host/subprocess_timeout.py", line 40, in subprocess_with_timeout    close_fds=True)  File "/usr/lib64/python2.6/subprocess.py", line 642, in __init__    errread, errwrite)  File "/usr/lib64/python2.6/subprocess.py", line 1234, in _execute_child    child_exception = pickle.loads(data)OSError: [Errno 2] No such file or directory



故障排查

这里设置的max user processes为65535已经非常大了,一般来说是达不到这个瓶颈的

# ulimit -acore file size          (blocks, -c) 0data seg size           (kbytes, -d) unlimitedscheduling priority             (-e) 0file size               (blocks, -f) unlimitedpending signals                 (-i) 127452max locked memory       (kbytes, -l) 64max memory size         (kbytes, -m) unlimitedopen files                      (-n) 65535pipe size            (512 bytes, -p) 8POSIX message queues     (bytes, -q) 819200real-time priority              (-r) 0stack size              (kbytes, -s) 10240cpu time               (seconds, -t) unlimitedmax user processes              (-u) 65535virtual memory          (kbytes, -v) unlimitedfile locks                      (-x) unlimited


现在系统的总进程数仅仅一百多个,我们要检查每个进程对应有多少个线程

# ps -ef|wc -l

169


已知这台服务器上主要跑的是java进程,所以重点查看java进程对应的线程数,找到30315这个进程对应约32110个线程,在加上其他进程和线程数,总数超过65535,NameNode无法在申请到多余的线程,所以报错

# pgrep java

1680

5482

19662

28770

30315

35902


# for i in `pgrep java`; do ps -T -p $i |wc -l; done

15

49

30

53

32110

114


# ps -T -p 30315|wc -l

32110


或者通过top -H 命令查看

# top -H

top - 10:44:58 up 779 days, 19:34, 3 users, load average: 0.01, 0.05, 0.05

Tasks: 32621 total, 1 running, 32620 sleeping, 0 stopped, 0 zombie

Cpu(s): 2.8%us, 4.1%sy, 0.0%ni, 93.1%id, 0.0%wa, 0.0%hi, 0.0%si, 0.0%st

Mem: 16334284k total, 15879392k used, 454892k free, 381132k buffers

Swap: 4194296k total, 0k used, 4194296k free, 8304400k cached


解决方法

找到了问题的原因,我们可以重新设定max user processes的值为100000,再次启动NameNode成功

#echo "100000" > /proc/sys/kernel/threads-max

#echo "100000" > /proc/sys/kernel/pid_max (默认32768)

#echo "200000" > /proc/sys/vm/max_map_count (默认65530)


#vim /etc/security/limits.d/90-nproc.conf

* soft nproc unlimited

root soft nproc unlimited


#vim /etc/security/limits.conf

* soft nofile 65535

* hard nofile 65535

* hard nproc 100000

* soft nproc 100000


# ulimit -u

100000


0