Pages

Thursday, September 23, 2010

VMware Virtual Infrastructure 3 (VI3) - Adding a Raw Device Mapping hard disk grayed out

We still have a client who are still on VI3 and I had a maintenance window today to reconfigure their virtualized SQL cluster and while executing my plan, I ran into an issue with mapping a newly provisioned LUN as an RDM. While this isn’t much of a difficult problem to solve, I’d like to blog it in case someone happens to Google this problem for a quick answer.

So here I am going through the regular tasks of attaching a new disk to an existing virtual machine.

image

image

…and I get to the screen with:

A virtual disk is composed of one or more files on the host file system. Together these files appear as a single hard disk to the guest operating system. Select the type of disk to use from the choices below.

Disk

Create a new virtual disk

Choose this option to create a new virtual disk.

Use an existing virtual disk

Choose this option to reuse a previously configured virtual disk.

Raw Device Mappings

Give your virtual machine direct access to the SAN. This option allows you to use existing SAN commands to manage the storage and continue to access it using a datastore.

image

The problem I had was that the option is grayed out. Fortunately, I sort of knew what might be the problem so I navigated to the Configuration tab, Storage Adapters and found that the newly provisioned LUN hasn’t shown up yet:

image

What I ended up doing was do a quick Rescan (top right hand corner in the screenshot above) of on the HBAs, got the LUN to show up and then went back to adding the disk. From there on, the option was available.

I hope this post will end up saving someone’s troubleshooting time.

Tuesday, September 21, 2010

Backup Exec 2010 fails with “e000848c - Unable to attach to a resource” and other errors

While configuring backups for a customer with Backup Exec 2010 to backup a vSphere 4 environment, we ran into multiple errors such as the following:

Final error: 0xe000848c - Unable to attach to a resource. Make sure that all selected resources exist and are online, and then try again. If the server or resource no longer exists, remove it from the selection list. Edit the selection list properties, click the View Selection Details tab, and then remove the resource.
Final error category: Resource Errors

For additional information regarding this error refer to link V-79-57344-33932

V-79-57344-38277 – Unable to open a disk of the virtual machine

VixDiskLib_Open() reported the error:
V-79-57344-38277 - Unable to open a disk of the virtual machine.

VixDiskLib_Open() reported the error:

image

Problem Resolution

After troubleshooting the problem for an hour and being pressed for time, we ended up calling Symantec since we had a support contract. What ended up being the problem was because we were not using the R2 version. Here’s the version we were initially using:

Symantec

Backup Exec 2010

Media Server: Version 13.0 Rev. 2896 (64-bit)

Administration Console: Version 13.0 Rev. 2896 (64-bit)

Desktop and Laptop Option: Version 3.1 Rev. 3.42.44a

image

Once we upgraded to the R2 version, the backups began to work:

Symantec

Backup Exec 2010 R2

Media Server: Version 13.0 Rev. 4164 (64-bit)

Administration Console: Version 13.0 Rev. 4164 (64-bit)

Desktop and Laptop Option: Version 3.1 Rev. 3.43.17a

image

Here are some additional copy and paste of the error logs:

Job name : test

Job type : Backup

Job status : Failed

Job log : C:\Program Files\Symantec\Backup Exec\Data\BEX_WCITORBTBEXEC_00155.xml

Server name : BEXEC

Selection list name : test-1
Device name : Test

Target name : Test

Media set name : Daily Full Backups
Error category : Resource Errors

Error : e000848c - Unable to attach to a resource. Make sure that all selected resources exist and are online, and then try again. If the server or resource no longer exists, remove it from the selection list. Edit the selection list properties, click the View Selection D

For additional information regarding this error refer to link V-79-57344-33932

-------------------------------------------------------------------------------------------------------------------------------------------------------------------

Set type : Backup

Set status : Completed

Set description : test
Resource name : \\VCVM\VMVCB::\\VCVM\VCGuestVm\vm\WFPS

Logon account : System Logon Account

Encryption used : None
Error : e0009585 - Unable to open a disk of the virtual machine.
Agent used : Yes

Advanced Open File Option used : No

-------------------------------------------------------------------------------------------------------------------------------------------------------------------

Job ended: September-15-10 at 3:47:55 PM
Completed status: Failed
Final error: 0xe000848c - Unable to attach to a resource. Make sure that all selected resources exist and are online, and then try again. If the server or resource no longer exists, remove it from the selection list. Edit the selection list properties, click the View Selection Details tab, and then remove the resource.
Final error category: Resource Errors

For additional information regarding this error refer to link V-79-57344-33932

-------------------------------------------------------------------------------------------------------------------------------------------------------------------

Click an error below to locate it in the job log

Backup- VMVCB::\\VCVM\VCGuestVm\vm\FPS

V-79-57344-38277 - Unable to open a disk of the virtual machine.

VixDiskLib_Open() reported the error:
V-79-57344-38277 - Unable to open a disk of the virtual machine.

VixDiskLib_Open() reported the error:

Problem powering on a virtual machine - “Cannot open the disk ‘/vmfs/volumes/unique identifier/virtualmachine/file.vmdk’

While this error can be easily solved by reviewing the events in VI Client, I figure I should blog it anyways in case anyone was searching for this string and had ran out of ideas.

I was doing some maintenance on a hosted environment a few days ago where we had to add new NetApp shelf to the existing FAS that hosted storage for a 3 node ESX cluster with 20 virtual machines currently used to host a medical tracking application. The setup was nothing fancy and small enough to be easily built with the NetApp providing SAN storage with SAS drives and an iSCSI SAN with SATA drives for backups.

The work done in the maintenance window was mainly ESX and storage so we only had my storage colleague and myself. The problem I ran into early was that since the environment was so locked down, I was unable to get to the ESX hosts directly with VI Client when I shutdown Virtual Center (It’s VI3). I ended up hooking my laptop to the iSCSI storage’s port to get to the hosts directly.

Once we finished provisioning the new storage and started booting back up the VMs, I received an error on the 2 virtual machines configured with MSCS and SQL clustering. The VMs had RDMs and disks located on the iSCSI vmfs store for backups. The first error I saw was:

Failed to power on virtualMachine: A general system error occurred:

image

Definitely not helpful at all. The next message was much better:

Message on virtualMachine: Cannot open the disk ‘/vmfs/volumes/unique identifier/virtualmachine/file.vmdk

image

At the bottom of the Events entries in the Events Details window, we can see the following message:

Message on virtualMachine on esxHost in datacenterName: Cannot open the disk ‘/vmfs/volumes/unique identifier/virtualmachine/file.vmdk or one of….

Reason: Device or resource busy.

image

Once I saw the 2nd error I knew immediately why the VM wasn’t powering on (it has lost access to some of the disks) so I replaced the network cable from the iSCSI storage controller back to the port that it was on. From there on, the server booted back up.

Thursday, September 16, 2010

Windows Server 2008 R2 64-bit cluster verification error with VMware virtual machines

Ran into a problem a few months ago while building a MSCS (Microsoft Clustering Services) while verifying the cluster nodes when I would continue to get the following message:

An error occurred while executing the test. There was an error verifying the firewall configuration. An item with the same key has already been added.

After searching for awhile for the answer, I found that this was because the NICs on the virtual machines had the same GUIDs and MSCS doesn’t like that. This generally wouldn’t be a problem if the servers were physical because the GUIDs would be different but this customer opted to deploy the cluster in their vSphere environment and since the servers were deployed from templates, the GUIDs ended up being the same. One of the blog posts I found indicated that I can remove the NIC from within Windows but when I tried doing so through device manager, the GUID ended up being the same. What I ended up doing to resolve the problem was:

  1. Uninstall the NIC from Device Manager within Windows.
  2. Shut off the virtual machine.
  3. Removed the NIC by editing the virtual machine’s settings from vCenter, click OK.
  4. Edit the virtual machine’s settings and add a new NIC, click OK.

Once Windows completed the startup, I went into the registry and can now see the new NIC haven’t a different GUID.

NetApp Virtual Storage Console for VMware vSphere Setup Wizard ended prematurely

Ran into an issue installing NetApp’s Virtual Storage Console for VMware vSphere the other day on a Windows Server 2008 R2 64-bit server where the install would fail with:

NetApp Virtual Storage Console for VMware vSphere Setup Wizard ended prematurely

NetApp Virtual Storage Console for VMware vSphere Setup Wizard ended prematurely because of an error. Your system has not been modified. To install this program at a later time…

Click the Finish button to exit the Setup Wizard.

image

For those who don’t want to read and view all the screenshots I’m including, the problem is because the install wasn’t executing with “run as administrator”. Another problem with the package I was using was that it was an .msi package and right clicking on it does not provide me with the option so what I ended up doing was opening up a command prompt as administrator, navigate to the folder with the netapp_vsc_1.0.msi package, then executing it from there. The following shows the process:

image

image

image

image

image

image

image

image

image

So even thought we got the User Account Control (UAC) prompt, the install will still eventually fail:

image

image

Failed.

image

To get the package to install properly, open the command prompt running as an administrator and then execute the netapp_vsc_1.0.msi package:

image

image

image

image

image

image

image

image

image

image

image

image

Make sure you add the IP of the server with the NetApp Virtual Storage Console to the trusted sites of the IE so the tab within vCenter is displayed properly.

image

image

image

image

image

image

image

image

image

image

image

image

Wednesday, September 15, 2010

Troubleshooting SQL Backup Job Failure – Logs appear to be truncated

I don’t do quite as much SQL DBA tasks these days as the projects I’ve been working on at the current company aren’t as involved with SQL as the previous company I was with. With that being said, there are a few environments with SQL clusters with issues that are escalated to me every so often and one of them began having backup jobs from maintenance plans that I’ve set up fail once every week. As I began troubleshooting and reviewing the logs, I noticed that Job log error had an error logged but the information was somewhat truncated. The following is the flow of troubleshooting that I went through to ultimately figuring out why the jobs were failing.

The first logs I went to look at were the Error Logs section:

image

Unfortunately, the logs didn’t really tell me much (I also had the SQL Server Agent logs selected but did not see any errors) so I went ahead and checked the SQL Server Agent Jobs logs and was able to see an error with a bit more information but it was truncated:

image

Here’s a dump of the text from the error. In order for it to be displayed problem, it needs to be copied to notepad with the size of the window bigger:

Date 9/8/2010 2:00:01 AM
Log Job History (Temp backup to X.Subplan_1)

Step ID 1
Server xxxSQLCLUSTER
Job Name Temp backup to X.Subplan_1
Step Name Subplan_1
Duration 02:12:13
Sql Severity 0
Sql Message ID 0
Operator Emailed
Operator Net sent
Operator Paged
Retries Attempted 0

Message
Executed as user: XXXX\svc_sqlclusteragent. ...531.0 for 64-bit Copyright (C) Microsoft Corp 1984-2005. All rights reserved. Started: 2:00:01 AM Progress: 2010-09-08 02:00:10.94 Source: {7505B5D7-1DFB-4550-B691-085284227C50} Executing query "DECLARE @Guid UNIQUEIDENTIFIER EXECUTE msdb..sp...".: 100% complete End Progress Progress: 2010-09-08 02:00:13.58 Source: Maintenance Cleanup Task Executing query "EXECUTE master.dbo.xp_delete_file 0,N'X:\',N'bak',...".: 100% complete End Progress Progress: 2010-09-08 02:00:16.17 Source: Back Up Database (Full) Executing query "EXECUTE master.dbo.xp_create_subdir N'X:\\XXXX.xxx...".: 11% complete End Progress Progress: 2010-09-08 02:00:16.17 Source: Back Up Database (Full) Executing query "EXECUTE master.dbo.xp_create_subdir N'X:\\XXXX.xxx...".: 22% complete End Progress Progress: 2010-09-08 02:00:16.17 Source: Back Up Database (Full) Executing query "EXECUTE master.dbo.xp_create_subdir N'X:\\XXXX.xxx...".: 33% complete End Progress Progress: 2010-09-08 02:00:16.17 Source: Back Up Database (Full) Executing query "EXECUTE master.dbo.xp_create_subdir N'X:\\XXXX.xxx...".: 44% complete End Progress Progress: 2010-09-08 02:00:16.17 Source: Back Up Database (Full) Executing query "EXECUTE master.dbo.xp_create_subdir N'X:\\XXXX.xxx...".: 55% complete End Progress Progress: 2010-09-08 02:00:16.17 Source: Back Up Database (Full) Executing query "EXECUTE master.dbo.xp_create_subdir N'X:\\XXXX.xxx...".: 66% complete End Progress Progress: 2010-09-08 02:00:16.17 Source: Back Up Database (Full) Executing query "EXECUTE master.dbo.xp_create_subdir N'X:\\master' ...".: 77% complete End Progress Progress: 2010-09-08 02:00:16.17 Source: Back Up Database (Full) Executing query "EXECUTE master.dbo.xp_create_subdir N'X:\\model' ".: 88% complete End Progress Progress: 2010-09-08 02:00:16.19 Source: Back Up Database (Full) Executing query "EXECUTE master.dbo.xp_create_subdir N'X:\\msdb' ".: 100% complete End Progress Progress: 2010-09-08 02:14:24.05 Source: Back Up Database (Full) Executing query "BACKUP DATABASE [XXXX.xxx.Production] TO DISK = N...".: 50% complete End Progress Progress: 2010-09-08 02:27:30.57 Source: Back Up Database (Full) Executing query "declare @backupSetId as int select @backupSetId =...".: 100% complete End Progress Progress: 2010-09-08 02:58:42.92 Source: Back Up Database (Full) Executing query "BACKUP DATABASE [XXXX.xxx.Production.DataStore] TO...".: 50% complete End Progress Progress: 2010-09-08 04:02:53.63 Source: Back Up Database (Full) Executing query "declare @backupSetId as int select @backupSetId =...".: 100% complete End Progress Progress: 2010-09-08 04:03:14.19 Source: Back Up Database (Full) Executing query "BACKUP DATABASE [XXXX.xxx.Production.MemberShip] T...".: 50% complete End Progress Progress: 2010-09-08 04:03:23.57 Source: Back Up Database (Full) Executing query "declare @backupSetId as int select @backupSetId =...".: 100% complete End Progress Progress: 2010-09-08 04:05:09.59 Source: Back Up Database (Full) Executing query "BACKUP DATABASE [XXXX.xxx.Training] TO DISK = N'X...".: 50% complete End Progress Progress: 2010-09-08 04:05:56.55 Source: Back Up Database (Full) Executing query "declare @backupSetId as int select @backupSetId =...".: 100% complete End Progress Progress: 2010-09-08 04:07:41.12 Source: Back Up Database (Full) Executing query "BACKUP DATABASE [XXXX.xxx.Training.DataStore] TO ...".: 50% complete End Progress Progress: 2010-09-08 04:08:27.17 Source: Back Up Database (Full) Executing query "declare @backupSetId as int select @backupSetId =...".: 100% complete End Progress Error: 2010-09-08 04:08:53.25 Code: 0xC00291EC Source: Back Up Database (F... The package execution fa... The step failed.

The error is indicated at the bottom where it reads:

Error: 2010-09-08 04:08:53.25 Code: 0xC00291EC Source: Back Up Database (F... The package execution fa... The step failed.

The information’s not very helpful with the truncated text so I went back to SQL Server Business Management Studio to look around. I ended up at the Maintenance Plan node’s actual job:

image

Drilling down to the actual job and clicking history gave me more information:

image

The error number is: -1073573396

The error message is: Failed to acquire connection "Local server connection". Connection may not be configured correctly or you may not have the right permissions on this connection.

From here I was able to review the time the job failed, what else was going on and noticed that there was an overlap during the time of the backups and database re-indexing which probably lead to the backup not able to obtain the lock it needed on the database. What I ended up doing was adjust the timing of the two jobs.