Sun Microsystems, Inc.
spacerspacer
spacer www.sun.com docs.sun.com |
spacer
black dot
 
 
2.  Error Messages Message IDs 600000-699999  Previous   Contents   Next 
   
 

 

657885 sigwait: %s

Description: The cl_apid was unable to configure its signal handling functionality, so it is unable to run.

Solution: Save a copy of the /var/adm/messages files on all nodes and contact your authorized Sun service provider for assistance in diagnosing and correcting the problem.

 

658329 CMM: Waiting for initial handshake to complete.

Description: The userland CMM has not been able to complete its initial handshake protocol with its counterparts on the other cluster nodes, and will only be able to join the cluster after this is completed.

Solution: This is an informational message, no user action is needed.

 

658555 Retrying to retrieve the resource information.

Description: An update to cluster configuration occurred while resource properties were being retrieved

Solution: Ignore the message.

 

659665 kill -KILL: %s

Description: The rpc.fed server is not able to stop a tag that timed out, and the error message is shown. An error message is output to syslog.

Solution: Save the /var/adm/messages file. Examine other syslog messages occurring around the same time on the same node, to see if the cause of the problem can be identified.

 

659827 CCR: Can't access CCR metadata on node %s errno = %d.

Description: The indicated error occurred when CCR is trying to access the CCR metadata on the indicated node. The errno value indicates the nature of the problem. errno values are defined in the file /usr/include/sys/errno.h. An errno value of 28(ENOSPC) indicates that the root files system on the node is full. Other values of errno can be returned when the root disk has failed(EIO).

Solution: There may be other related messages on the node where the failure occurred. These may help diagnose the problem. If the root file system is full on the node, then free up some space by removing unnecessary files. If the root disk on the afflicted node has failed, then it needs to be replaced. If the cluster repository is corrupted, boot the indicated node in -x mode to restore it from backup. The cluster repository is located at /etc/cluster/ccr/.

 

660332 launch_validate: fe_set_env_vars() failed for resource <%s>, resource group <%s>, method <%s>

Description: The rgmd was unable to set up environment variables for method execution, causing a VALIDATE method invocation to fail. This in turn will cause the failure of a creation or update operation on a resource or resource group.

Solution: Examine other syslog messages occurring at about the same time to see if the problem can be identified. Re-try the creation or update operation. If the problem recurs, save a copy of the /var/adm/messages files on all nodes and contact your authorized Sun service provider for assistance.

 

660368 CCR: CCR service not available, service is %s.

Description: The CCR service is not available due to the indicated failure.

Solution: Reboot the cluster. Also contact your authorized Sun service provider to determine whether a workaround or patch is available.

 

660974 file specified in USER_ENV %s does not exist

Description: 'User_env' property was set when configuring the resource. File specified in 'User_env' property does not exist or is not readable. File should be specified with fully qualified path.

Solution: Specify existing file with fully qualified file name when creating resource. If resource is already created, please update resource property 'User_env'.

 

660974 file specified in USER_ENV %s does not exist

Description: 'User_env' property was set when configuring the resource. File specified in 'User_env' property does not exist or is not readable. File should be specified with fully qualified path.

Solution: Specify existing file with fully qualified file name when creating resource. If resource is already created, please update resource property 'User_env'.

 

661084 liveCache was stopped by the user outside of Sun Cluster. Sun Cluster will suspend monitoring until liveCache is again started up successfully outside of Sun Cluster.

Description: When Sun Cluster tries to bring up liveCache, it detects that liveCache was brought down by user intendedly outside of Sun Cluster. Suu Cluster will not try to restart it under the control of Sun Cluster until liveCache is started up successfully again by the user. This behaviour is enforced across nodes in the cluster.

Solution: Informative message. No action is needed.

 

661560 All the SUNW.HAStoragePlus resources that this resource depends on are online on the local node. Proceeding with the checks for the existence and permissions of the start/stop/probe commands.

Description: The HAStoragePlus resource that this resource depends on is local to this node. Proceeding with the rest of the validation checks.

Solution: This message is informational; no user action is needed.

 

661560 All the SUNW.HAStoragePlus resources that this resource depends on are online on the local node. Proceeding with the checks for the existence and permissions of the start/stop/probe commands.

Description: This is an informational message which means that the SUNW.HAStoragePlus resource(s) that this application resource depends on is online on the local node and therefore the validation checks related to start/stop/probe commands will be carried out on the local node.

Solution: None.

 

661778 clcomm: memory low: freemem 0x%x

Description: The system is reporting that the system has a very low level of free memory.

Solution: If the system fails soon after this message, then there is a significantly greater chance that the system ran out of memory. In which case either install more memory or reduce system load. When the system continues to function, this means that the system recovered and no user action is required.

 

662056 Failed to shutdown lockd gracefully.

Description: Not available at this time.

Solution: Not available at this time.

 

662516 SIOCGLIFNUM: %s

Description: The ioctl command with this option failed in the cl_apid. This error may prevent the cl_apid from starting up.

Solution: Examine other syslog messages occurring at about the same time to see if the problem can be identified. Save a copy of the /var/adm/messages files on all nodes and contact your authorized Sun service provider for assistance in diagnosing and correcting the problem.

 

663089 clexecd: %s: sigwait returned %d. Exiting.

Description: clexecd program has encountered a failed sigwait(3C) system call. The error message indicates the error number for the failure.

Solution: The clexecd program will exit and the node will be halted or rebooted to prevent data corruption. Contact your authorized Sun service provider to determine whether a workaround or patch is available.

 

663293 reservation error(%s) - do_status() error for disk %s

Description: The device fencing program has encountered errors while trying to access a device. All retry attempts have failed.

Solution: The action which failed is a scsi-2 ioctl. These can fail if there are scsi-3 keys on the disk. To remove invalid scsi-3 keys from a device, use 'scdidadm -R' to repair the disk (see scdidadm man page for details). If there were no scsi-3 keys present on the device, then this error is indicative of a hardware problem, which should be resolved as soon as possible. Once the problem has been resolved, the following actions may be necessary: If the message specifies the 'node_join' transition, then this node may be unable to access the specified device. If the failure occurred during the 'release_shared_scsi2' transition, then a node which was joining the cluster may be unable to access the device. In either case, access can be reacquired by executing '/usr/cluster/lib/sc/run_reserve -c node_join' on all cluster nodes. If the failure occurred during the 'make_primary' transition, then a device group may have failed to start on this node. If the device group was started on another node, it may be moved to this node with the scswitch command. If the device group was not started, it may be started with the scswitch command. If the failure occurred during the 'primary_to_secondary' transition, then the shutdown or switchover of a device group may have failed. If so, the desired action may be retried.

 

663851 Failover %s data services must have exactly one value for extension property %s.

Description: Failover data services must have one and only one value for Confdir_list.

Solution: Create a failover resource group for each configuration file.

 

663851 Failover %s data services must have exactly one value for extension property %s.

Description: Failover data services must have one and only one value for Confdir_list.

Solution: Create a failover resource group for each configuration file.

 

663897 clcomm: Endpoint %p: %d is not an endpoint state

Description: The system maintains information about the state of an Endpoint. The Endpoint state is invalid.

Solution: Contact your authorized Sun service provider to determine whether a workaround or patch is available.

 

663943 Quorum: Unable to reset node information on quorum disk.

Description: This node was unable to reset some information on the quorum device. This will lead the node to believe that its partition has been preempted. This is an internal error. If a cluster gets divided into two or more disjoint subclusters, exactly one of these must survive as the operational cluster. The surviving cluster forces the other subclusters to abort by grabbing enough votes to grant it majority quorum. This is referred to as preemption of the losing subclusters.

Solution: Contact your authorized Sun service provider to determine whether a workaround or patch is available.

 
 
 
  Previous   Contents   Next