Sun Microsystems, Inc.
spacerspacer
spacer www.sun.com docs.sun.com |
spacer
black dot
 
 
2.  Error Messages Message IDs 500000-599999  Previous   Contents   Next 
   
 

 

523933 Although there are no other potential masters, RGM is failing resource group <%s> off of node <%d> because there are other current healthy masters.

Description: The resource group was brought OFFLINE on the node specified, probably because of a public network failure on that node. The operation was performed despite the lack of a healthy candidate node to host the resource group, because the resource group was currently mastered by at least one other healthy node.

Solution: No action required. If desired, examine other syslog messages on the node in question to determine the cause of the network failure.

 

525197 No network address resources in resource group.

Description: The cl_apid encountered an invalid property value. If it is trying to start, it will terminate. If it is trying to reload the properties, it will use the old properties instead.

Solution: Save a copy of the /var/adm/messages files on all nodes and contact your authorized Sun service provider for assistance in diagnosing and correcting the problem.

 

525628 CMM: Cluster has reached quorum.

Description: Enough nodes are operational to obtain a majority quorum; the cluster is now moving into operational state.

Solution: This is an informational message, no user action is needed.

 

526056 Resource <%s> of Resource Group <%s> failed pingpong check on node <%s>. The resource group will not be mastered by that node.

Description: A scha_control(1HA,3HA) call has failed because no healthy new master could be found for the resource group. A given node is considered unhealthy for a given resource if that same resource has recently initiated a failover off of that node by a previous scha_control call. In this context, "recently" means within the past Pingpong_interval seconds, where Pingpong_interval is a user-configurable property of the resource group. The default value of Pingpong_interval is 3600 seconds. This check is performed to avoid the situation where a resource group repeatedly "ping-pongs" or moves back and forth between two or more nodes, which might occur if some external problem prevents the resource group from running successfully on *any* node.

Solution: A properly-implemented resource monitor, upon encountering the failure of a scha_control call, should sleep for awhile and restart its probes. If the resource remains unhealthy, the problem that caused the scha_control call to fail (such as pingpong check described above) will eventually resolve, permitting a later scha_control request to succeed. Therefore, no user action is required. If the system administrator wishes to permit failovers to be attempted even at the risk of ping-pong behavior, the Pingpong_interval property of the resource group should be set to a smaller value.

 

526403 ff_open: %s

Description: A server (rpc.pmfd or rpc.fed) was not able to establish a link to the failfast device, which ensures that the host aborts if the server dies. The error message is shown. An error message is output to syslog.

Solution: Save the /var/adm/messages file. Contact your authorized Sun service provider to determine whether a workaround or patch is available.

 

526492 Service object [%s, %s, %d] removed from group '%s'

Description: A specific service known by its unique name SAP (service access point), the three-tuple, has been deleted in the designated group.

Solution: This is an informational message, no user action is needed.

 

526846 Daemon <%s> is not running.

Description: The HA-NFS fault monitor detected that the specified daemon is no longer running.

Solution: No action. The fault monitor would restart the daemon. If it doesn't happen, reboot the node.

 

527210 Unable to read %s: %s.

Description: Not available at this time.

Solution: Not available at this time.

 

527795 clexecd: setrlimit returned %d

Description: clexecd program has encountered a failed setrlimit() system call. The error message indicates the error number for the failure.

Solution: Contact your authorized Sun service provider to determine whether a workaround or patch is available.

 

528020 CCR: Remove table %s failed.

Description: The CCR failed to remove the indicated table.

Solution: The failure can happen due to many reasons, for some of which no user action is required because the CCR client in that case will handle the failure. The cases for which user action is required depends on other messages from CCR on the node, and include: If it failed because the cluster lost quorum, reboot the cluster. If the root file system is full on the node, then free up some space by removing unnecessary files. If the root disk on the afflicted node has failed, then it needs to be replaced. If the cluster repository is corrupted as indicated by other CCR messages, then boot the offending node(s) in -x mode to restore the cluster repository from backup. The cluster repository is located at /etc/cluster/ccr/.

 

528499 scsblconfig not configured correctly.

Description: The specified file has not been configured correctly, or it does not have all the required settings.

Solution: Please verify that required variables (according to the installation instructions for this data service) are correctly configured in this file. Try to manually source this file in korn shell (". scsblconfig"), and verify if the required variables are getting set correctly.

 

528566 Method <%s> on resource <%s>, resource group <%s>, is_frozen=<%d>: Method timed out.

Description: A method execution has exceeded its configured timeout and was killed by the rgmd. Depending on which method was being invoked and the Failover_mode setting on the resource, this might cause the resource group to fail over or move to an error state.

Solution: Consult resource type documentation to diagnose the cause of the method failure. Other syslog messages occurring just before this one might indicate the reason for the failure. After correcting the problem that caused the method to fail, the operator may choose to issue an scswitch(1M) command to bring resource groups onto desired primaries. Note, if the indicated value of is_frozen is 1, this might indicate an internal error in the rgmd. Please save a copy of the /var/adm/messages files on all nodes, and report the problem to your authorized Sun service provider.

 

529131 Method <%s> on resource <%s>: RPC connection error.

Description: An attempted method execution failed, due to an RPC connection problem. This failure is considered a method failure. Depending on which method was being invoked and the Failover_mode setting on the resource, this might cause the resource group to fail over or move to an error state; or it might cause an attempted edit of a resource group or its resources to fail.

Solution: Examine other syslog messages occurring around the same time on the same node, to see if the cause of the problem can be identified. If the same error recurs, you might have to reboot the affected node. After the problem is corrected, the operator may choose to issue an scswitch(1M) command to bring resource groups onto desired primaries, or re-try the resource group update operation.

 

529191 clexecd: Sending fd to workerd returned %d. Exiting.

Description: There was some error in setting up interprocess communication in the clexecd program.

Solution: Contact your authorized Sun service provider to determine whether a workaround or patch is available.

 

529407 resource group %s state on node %s change to %s

Description: This is a notification from the rgmd that a resource group's state has changed. This may be used by system monitoring tools.

Solution: This is an informational message, no user action is needed.

 

530064 reservation error(%s) - do_enfailfast() error for disk %s

Description: The device fencing program has encountered errors while trying to access a device. All retry attempts have failed.

Solution: This may be indicative of a hardware problem, which should be resolved as soon as possible. Once the problem has been resolved, the following actions may be necessary: If the message specifies the 'node_join' transition, then this node may be unable to access the specified device. If the failure occurred during the 'release_shared_scsi2' transition, then a node which was joining the cluster may be unable to access the device. In either case, access can be reacquired by executing '/usr/cluster/lib/sc/run_reserve -c node_join' on all cluster nodes. If the failure occurred during the 'make_primary' transition, then a device group may have failed to start on this node. If the device group was started on another node, it may be moved to this node with the scswitch command. If the device group was not started, it may be started with the scswitch command. If the failure occurred during the 'primary_to_secondary' transition, then the shutdown or switchover of a device group may have failed. If so, the desired action may be retried.

 

530492 fatal: ucmm_initialize() failed

Description: The daemon indicated in the message tag (rgmd or ucmmd) was unable to initialize its interface to the low-level cluster membership monitor. This is a fatal error, and causes the node to be halted or rebooted to avoid data corruption. The daemon produces a core file before exiting.

Solution: Save a copy of the /var/adm/messages files on all nodes, and of the core file generated by the daemon. Contact your authorized Sun service provider for assistance in diagnosing the problem.

 

530603 Warning: Scalable service group for resource %s has already been created.

Description: It was not expected that the scalable services group for the named resource existed.

Solution: Rebooting all nodes of the cluster will cause the scalable services group to be deleted.

 

530828 Failed to disconnect from host %s and port %d.

Description: The data service fault monitor probe was trying to disconnect from the specified host/port and failed. The problem may be due to an overloaded system or other problems. If such failure is repeated, Sun Cluster will attempt to correct the situation by either doing a restart or a failover of the data service.

Solution: If this problem is due to an overloaded system, you may consider increasing the Probe_timeout property.

 

530938 Starting NFS daemon %s.

Description: The specified NFS daemon is being started by the HA-NFS implementation.

Solution: This is an informational message. No action is needed.

 

531148 fatal: thr_create stack allocation failure: %s (UNIX error %d)

Description: The rgmd was unable to create a thread stack, most likely because the system has run out of swap space. The rgmd will produce a core file and will force the node to halt or reboot to avoid the possibility of data corruption.

Solution: Rebooting the node has probably cured the problem. If the problem recurs, you might need to increase swap space by configuring additional swap devices. See swap(1M) for more information.

 
 
 
  Previous   Contents   Next