Nagios是一款企業級開源軟體,專注于監控伺服器上服務是否正常,不生成圖形,提供報警機制,郵件或者短信發送監控狀态,它通過各種插件實作不同的功能。
Nagios 監控平台主程式
Nagios-plugins 必選插件
NRPE 監控遠端伺服器的主機資源
NSClient++ 用于監控Windows主機
NDOUtils 将資料寫入資料庫
執行個體應用:
1 監控快速部署
監控需要安裝http php nagios nagios-plugins NRPE軟體包
yum install -y gd gd-devel openssl openssl-devel httpd php gcc glibc glib-common make wget
net-snmp
setenforce 0
iptables -F
安裝nagios 源碼包下載下傳安裝
wget http://sourceforge.net/projects/nagios/files/nagios-3.x/nagios-3.5.0/nagios-3.5.0.tar.gz/download
groupadd nagios
useradd -g nagios nagios
tar -zxf nagios-3.5.0.tar.gz -C /usr/src/
cd /usr/src/nagios
./configure --with-nagios-user=nagios --with-nagios-group=nagios
make all
make install
make install-init #安裝啟動腳本
make install-commandmode #安裝與配置目錄權限
make install-config #安裝配置檔案模闆
make install-webconf #web監控界面配置
安裝nagios-plugins和nrpe
wget http://nchc.dl.sourceforge.net/project/nagiosplug/nagiosplug/1.4.16/nagios-plugins-1.4.16.tar.gz
tar -zxf nagios-plugins-1.4.16.tar.gz -C /usr/src/
cd /usr/src/nagios-plugins-1.4.16
./configure --prefix=/usr/local/nagios/
make && make install
wget wget http://nchc.dl.sourceforge.net/project/nagios/nrpe-2.x/nrpe-2.14/nrpe-2.14.tar.gz
tar -zxf nrpe-2.14.tar.gz -C /usr/src/
cd /usr/src/nrpe-2.14
./configure
make install-plugin
make install-daemon
make install-daemon-config
chown -R nagions.nagions /usr/local/nagios
建立賬戶資訊
htpasswd -c /usr/local/nagions/etc/htpasswd.users tomcat
iptables -I INPUT -p tcp --dport 80 -j ACCEPT
service iptables save
啟動服務
service httpd start
/etc/init.d/nagios start
chkconfig httpd on
chkconfig --add nagios
chkconfig nagios on
2 修改配置檔案
nagios的配置檔案較多,主要位于/usr/local/nagios/etc 下
nagios.conf 主配置檔案
nrpe.cfg 遠端監控配置檔案
cgi.conf CGI配置檔案
commands.cfg 指令定義檔案
contacts.cfg 定義聯系人檔案
timepreriods.cfg 時間周期定義檔案
tempaltes.cfg 對象定義參考模闆
localhost.cfg 監控本機配置模闆
printer.cfg 監控列印機模闆
switch.cfg 監控交換模闆
windows.cfg 監控Windows配置模闆
很多配置檔案無需修改可以直接使用
修改主配置檔案nagios.cfg,主要是用cfg_file配置加載其他配置檔案。
vim /usr/local/nagios/etc/nagios.cfg
cfg_file=/usr/local/nagios/etc/objects/commands.cfg
cfg_file=/usr/local/nagios/etc/objects/contacts.cfg
cfg_file=/usr/local/nagios/etc/objects/templates.cfg
cfg_file=/usr/local/nagios/etc/objects/timeperiods.cfg
cfg_file=/usr/local/nagios/etc/objects/localhost.cfg
cfg_file=/usr/local/nagios/etc/web1.cfg
cfg_file=/usr/local/nagios/etc/web2.cfg
修改CGI配置檔案cgi.cfg,添加tomcat賬戶進來
vim /usr/local/nagios/etc/cgi.cfg
default_user_name=tomcat
authorized_for_system_information=nagiosadmin,tomcat
authorized_for_configuration_information=nagiosadmin,tomcat
authorized_for_system_commands=nagiosadmin,tomcat
authorized_for_all_services=nagiosadmin,tomcat
authorized_for_all_hosts=nagiosadmin,tomcat
authorized_for_all_service_commands=nagiosadmin,tomcat
authorized_for_all_host_commands=nagiosadmin,tomcat
修改指令配置檔案command.cfg,定義指令實作的方式,如郵件報警,使用工具,内容格式等。
vim /usr/local/nagios/etc/objects/commands.cfg
define command{
command_name check_nrpe
command_line $USER1$/check_nrpe -H $HOSTADDRESS$ -t 30 -c $ARG1$
}
command_name check_nrpe_args
command_line $USER1$/check_nrpe -H $HOSTADDRESS$ -t 30 -c $ARG1$ -a $ARG2$
修改聯系人配置檔案contacts.cfg 報警的聯系人及聯系方式
define contact{
contact_name nagiosadmin
use generic-contact
alias Nagios Admin
email [email protected]
修改報警時間周期timeperiods.cfg
vim /usr/local/nagios/etc/objects/timeperiods.cfg
define timeperiods{
timeperiod_name 24x7 #監控所有時間段(7*24小時)
alias 24 Hours A Day, 7 Days A Week
sunday 00:00-24:00
monday 00:00-24:00
tuesday 00:00-24:00
wednesday 00:00-24:00
thursday 00:00-24:00
friday 00:00-24:00
saturday 00:00-24:00
修改本機的配置localhost.cfg
define host{
use linux-server
host_name duangr-1
alias duangr-1
address 192.168.56.10
define service{
use local-service
host_name duangr-1
service_description Host Alive
check_command check-host-alive
service_description Users
check_command check_local_users!20!50
service_description CPU
check_command check_local_load!5.0,4.0,3.0!10.0,6.0,4.0
service_description Disk Root
check_command check_local_disk!20%!10%!/
service_description Disk Home
check_command check_local_disk!20%!10%!/export/home
service_description Zombie Procs
check_command check_local_procs!5!10!Z
service_description Total Procs
check_command check_local_procs!250!400!RSZDT
service_description Swap Usage
check_command check_local_swap!20!10
修改模闆檔案templates.cfg
vi /usr/local/nagios/etc/objects/templates.cfg
#聯系人模闆generic-contact
name generic-contact
service_notification_period 24x7
host_notification_period 24x7
service_notification_options w,u,c,r,f,s
host_notification_options d,u,r,f,s
service_notification_commands notify-service-by-email
host_notification_commands notify-host-by-email
register 0
}
#定義generic-host主機模闆
name generic-host
notifications_enabled 1
event_handler_enabled 1
flap_detection_enabled 1
failure_prediction_enabled 1
process_perf_data 1
retain_status_information 1
retain_nonstatus_information 1
notification_period 24x7
register 0
#定義Linux主機模闆
name linux-server
use generic-host
check_period 24x7
check_interval 5
retry_interval 1
max_check_attempts 10
check_command check-host-alive
notification_period workhours
notification_interval 120
notification_options d,u,r
contact_groups admins
register 0
建立遠端監控web1.cfg
vim /usr/local/nagios/etc/web1.cfg
define host{
use linux-server
host_name duangr-2
alias duangr-2
address 192.168.56.11
}
define service{
use local-service
host_name duangr-2
service_description Host Alive
check_command check-host-alive
service_description Users
check_command check_nrpe_args!check_users!5 10
service_description CPU
check_command check_nrpe_args!check_load!15,10,5 30,25,20
service_description Disk Root
check_command check_nrpe_args!check_disk!20% 10% /
service_description Disk /export/home
check_command check_nrpe_args!check_disk!20% 10% /export/home
use local-service
host_name duangr-2
service_description Procs Zombie
check_command check_nrpe_args!check_procs!5 10 Z
}
service_description Procs Total
check_command check_nrpe_args!check_procs_args!"-w400 -c600" }
service_description Swap Usage
check_command check_nrpe_args!check_swap!20% 10%
;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;
;; 下面是一些常用程序的監控,主要是雲平台相關程序
;; 監控crond程序
service_description PS: crond
check_command check_nrpe_args!check_procs_args!"-c1:1 -Ccrond" }
;; 監控zookeeper程序
service_description PS: QuorumPeerMain
check_command check_nrpe_args!check_procs_args!"-c1:1 -Cjava -aserver.quorum.QuorumPeerMain" }
;;監控storm的從節點程序
service_description PS: supervisor
check_command check_nrpe_args!check_procs_args!"-c1:1 -Cjava -adaemon.supervisor" }
;; 監控storm的主節點程序
service_description PS: nimbus
check_command check_nrpe_args!check_procs_args!"-c1:1 -Cjava -adaemon.nimbus" }
;; 監控MetaQ程序
service_description PS: MetaQ
check_command check_nrpe_args!check_procs_args!"-c1:1 -Cjava -ametamorphosis-server-w" }
;; 監控Redis程序
service_description PS: redis-server
check_command check_nrpe_args!check_procs_args!"-c1:1 -Credis-server" }
;; 監控hadoop主節點NameNode程序
service_description PS: NameNode
check_command check_nrpe_args!check_procs_args!"-c1:1 -Cjava -aserver.namenode.NameNode" }
;; 監控hadoop主節點SecondaryNameNode程序
service_description PS: SecondaryNameNode
check_command check_nrpe_args!check_procs_args!"-c1:1 -Cjava -aserver.namenode.SecondaryNameNode" }
;; 監控hadoop主節點ResourceManager程序
service_description PS: ResourceManager
check_command check_nrpe_args!check_procs_args!"-c1:1 -Cjava -aserver.resourcemanager.ResourceManager" }
;; 監控hadoop從節點DataNode程序
service_description PS: DataNode
check_command check_nrpe_args!check_procs_args!"-c1:1 -Cjava -aserver.datanode.DataNode" }
;;監控hadoop從節點NodeManager程序
service_description PS: NodeManager
check_command check_nrpe_args!check_procs_args!"-c1:1 -Cjava -aserver.nodemanager.NodeManager" }
由于duangr-2是遠端主機,是以使用check_nrpe_args指令來監控.
/etc/init.d/nagios restart
快速定位配置檔案問題所在指令
/usr/local/nagios/bin/nagios -V /usr/local/nagios/etc/nagios.cfg
3 被監控機安裝軟體 nagios-plugin nrpe
yum install -y openssl openssl-devel
useradd -g nagios -s /sbin/nologin nagios
tar -zxf nagios-plugins-2.1.6.tar.gz -C /usr/src/
cd /usr/src/nagios-plugins-2.1.6
./configure --prefix=/usr/local/nagios/ --with-nagios-user=nagios --with-nagios-group=nagios
修改用戶端的NRPE配置檔案
command[check_users]=/usr/local/nagios/libexec/check_users -w 5 -c 10
command[check_load]=/usr/local/nagios/libexec/check_load -w 15,10,5 -c 30,25,20
command[check_sda2]=/usr/local/nagios/libexec/check_disk -w 20% -c 10% -p /dev/sda2
command[check_swap]=/usr/local/nagios/libexec/check_disk -w 20% -c 10% -p /dev/shm
command[check_home]=/usr/local/nagios/libexec/check_disk -w 20% -c 10% -p /dev/mapper/VolGroup00-LogVol00
command[check_zombie_procs]=/usr/local/nagios/libexec/check_procs -w 5 -c 10 -s Z
command[check_total_procs]=/usr/local/nagios/libexec/check_procs -w 200 -c 300
command[check_ping81]=/usr/local/nagios/libexec/check_ping -H 10.155.0.1 -w 100.0,20% -c 500.0,60%#
command[check_hda1]=/usr/local/nagios/libexec/check_disk -w 20 -c 10 -p /dev/hda1
/usr/local/nagios/bin/nrpe -c /usr/local/nagios/etc/nrpe.cfg -d
echo "/usr/local/nagios/bin/nrpe -c /usr/local/nagios/etc/nrpe.cfg -d" >> /etc/rc.local
netstat -lnupt |grep 5666
iptables -I INPUT -p tcp --dport 5666 -j ACCEPT
檢查監控指令配置是否ok
/usr/local/nagios/libexec/check_nrpe -H localhost -c check_users -a 5 10
/usr/local/nagios/libexec/check_nrpe -H localhost -c check_load -a 15,10,5 30,25,20
/usr/local/nagios/libexec/check_nrpe -H localhost -c check_disk -a 20% 10% /
/usr/local/nagios/libexec/check_nrpe -H localhost -c check_procs -a 200 400 RSZDT
/usr/local/nagios/libexec/check_nrpe -H localhost -c check_swap -a 20% 10%
沒有問題就可以用浏覽器通路nagios了
本文轉自super李導51CTO部落格,原文連結:http://blog.51cto.com/superleedo/1891718 ,如需轉載請自行聯系原作者