Linux-HA Hardware Installation Guideline

This document (c) 1999 Volker Wiegand <Volker.Wiegand@suse.de>

This document serves as the starting point to plan, execute, and verify your hardware setup for a High Availability (HA) environment.

Contents

  1. Introduction
  2. Hardware Requirements
    1. Minimum Installation
    2. More Advanced Installation
    3. Fully Redundant Installation
  3. Hardware Setup and Test
    1. Serial Ports
    2. LAN Interfaces
    3. Other Devices
  4. Troubleshooting
  5. References

Introduction

With the high stability Linux has reached, this Operating System is well suited to be used for HA purposes. The Linux-HA project, based upon Harald Milz's HOWTO and Alan Robertson's Heartbeat code, provides the building blocks for a professional solution.

This document provides some advice on the initial planning, the installation and cans nterface eth1
heartbeat: 2000/01/18_14:26:46 notice: Using watchdog device: /dev/watchdog
heartbeat: 2000/01/18_14:26:46 error: Cannot open /proc/ha/.control: No such file or directory
heartbeat: 2000/01/18_14:26:56 warn: node linuxha2.linux-ha.org: is dead
heartbeat: 2000/01/18_14:26:56 INFO: Running /etc/ha.d/rc.d/status status
heartbeat: 2000/01/18_14:26:57 info: Requesting our resources.
heartbeat: 2000/01/18_14:26:58 INFO: Running /etc/ha.d/resource.d/IPaddr 192.168.85.3 status
heartbeat: 2000/01/18_14:26:58 INFO: Running /etc/ha.d/rc.d/ip-request ip-request
heartbeat: 2000/01/18_14:27:00 info: node linuxha2.linux-ha.org: status up
heartbeat: 2000/01/18_14:27:00 INFO: Running /etc/ha.d/rc.d/status status
heartbeat: 2000/01/18_14:27:28 Acquiring resource group: linuxha1.linux-ha.org 192.168.85.3 httpd smb mirror
heartbeat: 2000/01/18_14:27:28 INFO: Running /etc/ha.d/resource.d/mirror  start
heartbeat: 2000/01/18_14:27:29 INFO: Running /etc/rc.d/init.d/smb  start
heartbeat: 2000/01/18_14:27:30 INFO: Running /etc/rc.d/init.d/httpd  start
heartbeat: 2000/01/18_14:27:31 INFO: Running /etc/ha.d/resource.d/IPaddr 192.168.85.3 start
heartbeat: 2000/01/18_14:27:32 INFO: ifconfig eth0:0 192.168.85.3 netmask 255.255.255.0 broadcast 192.168.85.255
heartbeat: 2000/01/18_14:27:32 Sending Gratuitous Arp for 192.168.85.3 on eth0:0 [eth0]
NOTE:  Your log may differ depending on when you started heartbeat on linuxha2!!!  I waited just over 10 seconds.


OK, now try to ping your cluster's IP (192.168.85.3 in the example). If this works, telnet to it and verify you're on linuxha1.
Next, make sure your services are tied to the .3 address.  Bring up netscape and type in 192.168.85.3 for the URL.  For Samba, try to map the drive "\\192.168.85.3\test"  assuming you set up a share called "test".  See Samba docs to get that going.  As an aside, however, you'll want to use the "netbios name" parameter to have your Samba share listed under the cluster name and not the hostname of your cluster member!

NOTE: If you can't bring up the service IP address and you get ha-log entries similar to this:

        SIOCSIFADDR: No such device
        SIOCSIFFLAGS: No such device
        SIOCSIFNETMASK: No such device
        SIOCSIFBRDADDR: No such device
        SIOCSIFFLAGS: No such device
        SIOCADDRT: No such device
It may mean that you need to enable IP aliasing in your kernel build.  Check /usr/src/linux/.config for "CONFIG_IP_ALIAS=y" if you don't have it, you'll have the line "CONFIG_IP_ALIAS is not set".  Rebuild your kernel with IP aliasing enabled.
If this all works, you've got availability.  Now let's see if we have High Availability :-)

Take down linuxha1.  Kill power, kill heartbeat, whatever you have the stomach for, but don't just yank both the serial and eth1 heartbeat cables.  If you do that, you'll have services running on both nodes and when you re-connect the heartbeat, a bit of chaos....
Now ping the cluster IP. Approximately 5-10 seconds later it should start responding again. Telnet again and verify you're on linuxha2.  If it happens but takes more like 30 seconds, something is wrong.

If you get this far, it's probably working, but you should probably check all your heartbeats, too.
First, check your serial heartbeat.  Unplug the crossover cable from your eth1 NIC that you're using for your udp heartbeat.  Wait about 10 seconds.
Now, look at /var/log/ha-log on linuxha2 and make sure there's no line like this:
    1999/08/16_12:40:58 node linuxha1.linux-ha.org: is dead
If you get that, your serial heartbeat isn't working and your second node is taking over.  To avoid any problems, shut down heartbeat on the first node, then test your null modem cable.  Run the above serial tests again.

If your log is clean, great.  Re-connect the crossover cable.  Once that's done, disconnect the serial cable, wait 10 seconds and check the linuxha2 log again.
If it's clean, congrats!  If not, you can check /var/log/ha-log and /var/log/ha-debug for more clues.
 

Appendix A - Crossover Cable Construction

Your cable diagram should be as follows:

    Connector A     Connector B
 
 
Connector A Connector B
Pin # Pin #
1 3
2 6
3 1
6 2
4 7
5 8
7 4
8 5

Rev 1.1.0
(c) 2000 Rudy Pawul
rpawul@iso-ne.com ./usr/share/doc/heartbeat/HardwareGuide.html0100644000000000000000000003700507551103554017744 0ustar rootroot Preliminary Linux HA Hardware Installation Guide

Linux-HA Hardware Installation Guideline

This document (c) 1999 Volker Wiegand <Volker.Wiegand@suse.de>

This document serves as the starting point to plan, execute, and verify your hardware setup for a High Availability (HA) environment.

Contents

  1. Introduction
  2. Hardware Requirements
    1. Minimum Installation
    2. More Advanced Installation
    3. Fully Redundant Installation
  3. Hardware Setup and Test
    1. Serial Ports
    2. LAN Interfaces
    3. Other Devices
  4. Troubleshooting
  5. References

Introduction

With the high stability Linux has reached, this Operating System is well suited to be used for HA purposes. The Linux-HA project, based upon Harald Milz's HOWTO and Alan Robertson's Heartbeat code, provides the building blocks for a professional solution.

This document provides some advice on the initial planning, the installation and cans nterface eth1
heartbeat: 2000/01/18_14:26:46 notice: Using watchdog device: /dev/watchdog
heartbeat: 2000/01/18_14:26:46 error: Cannot open /proc/ha/.control: No such file or directory
heartbeat: 2000/01/18_14:26:56 warn: node linuxha2.linux-ha.org: is dead
heartbeat: 2000/01/18_14:26:56 INFO: Running /etc/ha.d/rc.d/status status
heartbeat: 2000/01/18_14:26:57 info: Requesting our resources.
heartbeat: 2000/01/18_14:26:58 INFO: Running /etc/ha.d/resource.d/IPaddr 192.168.85.3 status
heartbeat: 2000/01/18_14:26:58 INFO: Running /etc/ha.d/rc.d/ip-request ip-request
heartbeat: 2000/01/18_14:27:00 info: node linuxha2.linux-ha.org: status up
heartbeat: 2000/01/18_14:27:00 INFO: Running /etc/ha.d/rc.d/status status
heartbeat: 2000/01/18_14:27:28 Acquiring resource group: linuxha1.linux-ha.org 192.168.85.3 httpd smb mirror
heartbeat: 2000/01/18_14:27:28 INFO: Running /etc/ha.d/resource.d/mirror  start
heartbeat: 2000/01/18_14:27:29 INFO: Running /etc/rc.d/init.d/smb  start
heartbeat: 2000/01/18_14:27:30 INFO: Running /etc/rc.d/init.d/httpd  start
heartbeat: 2000/01/18_14:27:31 INFO: Running /etc/ha.d/resource.d/IPaddr 192.168.85.3 start
heartbeat: 2000/01/18_14:27:32 INFO: ifconfig eth0:0 192.168.85.3 netmask 255.255.255.0 broadcast 192.168.85.255
heartbeat: 2000/01/18_14:27:32 Sending Gratuitous Arp for 192.168.85.3 on eth0:0 [eth0]
NOTE:  Your log may differ depending on when you started heartbeat on linuxha2!!!  I waited just over 10 seconds.


OK, now try to ping your cluster's IP (192.168.85.3 in the example). If this works, telnet to it and verify you're on linuxha1.
Next, make sure your services are tied to the .3 address.  Bring up netscape and type in 192.168.85.3 for the URL.  For Samba, try to map the drive "\\192.168.85.3\test"  assuming you set up a share called "test".  See Samba docs to get that going.  As an aside, however, you'll want to use the "netbios name" parameter to have your Samba share listed under the cluster name and not the hostname of your cluster member!

NOTE: If you can't bring up the service IP address and you get ha-log entries similar to this:

        SIOCSIFADDR: No such device
        SIOCSIFFLAGS: No such device
        SIOCSIFNETMASK: No such device
        SIOCSIFBRDADDR: No such device
        SIOCSIFFLAGS: No such device
        SIOCADDRT: No such device
It may mean that you need to enable IP aliasing in your kernel build.  Check /usr/src/linux/.config for "CONFIG_IP_ALIAS=y" if you don't have it, you'll have the line "CONFIG_IP_ALIAS is not set".  Rebuild your kernel with IP aliasing enabled.
If this all works, you've got availability.  Now let's see if we have High Availability :-)

Take down linuxha1.  Kill power, kill heartbeat, whatever you have the stomach for, but don't just yank both the serial and eth1 heartbeat cables.  If you do that, you'll have services running on both nodes and when you re-connect the heartbeat, a bit of chaos....
Now ping the cluster IP. Approximately 5-10 seconds later it should start responding again. Telnet again and verify you're on linuxha2.  If it happens but takes more like 30 seconds, something is wrong.

If you get this far, it's probably working, but you should probably check all your heartbeats, too.
First, check your serial heartbeat.  Unplug the crossover cable from your eth1 NIC that you're using for your udp heartbeat.  Wait about 10 seconds.
Now, look at /var/log/ha-log on linuxha2 and make sure there's no line like this:
    1999/08/16_12:40:58 node linuxha1.linux-ha.org: is dead
If you get that, your serial heartbeat isn't working and your second node is taking over.  To avoid any problems, shut down heartbeat on the first node, then test your null modem cable.  Run the above serial tests again.

If your log is clean, great.  Re-connect the crossover cable.  Once that's done, disconnect the serial cable, wait 10 seconds and check the linuxha2 log again.
If it's clean, congrats!  If not, you can check /var/log/ha-log and /var/log/ha-debug for more clues.
 

Appendix A - Crossover Cable Construction

Your cable diagram should be as follows:

    Connector A     Connector B
 
 
Connector A Connector B
Pin # Pin #
1 3
2 6
3 1
6 2
4 7
5 8
7 4
8 5

Rev 1.1.0
(c) 2000 Rudy Pawul
rpawul@iso-ne.com ./u