Tom Limoncelli has a post up today entitled System administration needs more PhDs.
He makes some great observations and brings up a lot of interesting questions. The one that I think the others flow from is "Why are good practices so rarely adopted?"
My opinion, gained through observation, is that sysadmins arise from one of two places. Either they start out in relative isolation, or they come from an environment with multiple systems administrators.
The former develop their own ways of doing things through trial and error and/or research. This leads to endless ways of accomplishing the same or similar tasks. The utter heterogeneity of possible platform combinations lends itself to having each admin reinvent the wheel.
The latter typically have an established infrastructure in place, a well defined set of hardware, and a much more rigid structure of procedures and usually a bona fide methodology for change management.
The reason that the standalone sysadmin almost never resembles the well trained sysadmin is because best practices all seem to be vendor driven, reliant on a subset of devices and situations, and are hidden as well as possible behind support and contract agreements.
Those are hurdles the lone sysadmin faces AFTER he has discovered the "optimal solution", whatever that is. You mention puppet. Should you use cfengine or puppet? Unless you know about puppet, you'll use cfengine, unless you haven't heard of that either, in which case you'll roll your own. In my experience, you'll find $betterSolution right as you're implementing $bestSolutionYouKnowAbout.
I don't know whether there are more sysadmins in a single environment than in a plurality, but there are a _lot_ of sysadmins out there by themselves.
By themselves, sysadmins rely on their own cleverness, but together you get a synergy of ideas. The whole becomes smarter than the sum of the individuals, but most sysadmins never get to experience that. That's one of the reasons I started my blog. To shed light on what other people are doing, how they operate in their organizations, and so on.
Your books are a great resource for sysadmins, but the lone sysadmins of the world need to start communicating between themselves, and with the "institutional" admins out there. The same solution won't always work, but the sharing will go a way toward a meritocracy of
ideas.
Friday, October 31, 2008
Thursday, October 30, 2008
Dell's DRAC card sucks
I've worked with the my 1855 blade enclosure for a while, now, and I feel pretty confident in saying the following:
A) Dell's DRAC is a very useful device which facilitates remote administration
B) at least it would, if it didn't suck so much
The blade enclosure comes with a DRAC module that is inserted into a slot in the back. It's paired with an avocent KVM module connected to the blade units. You access the KVM through the DRAC web site, which is the real problem.
Every Dell technician I've complained to has said the same thing. "Yes, the DRAC is slow. Very slow, and underpowered". It's not just slow, it vacillates between borderline and completely unusable. On a good day, expect 3 minutes for the page to load. On a bad day, don't expect to load the whole page.

It also seems to occasionally lose track of the KVM. It's happened a few times so far, and there doesn't seem to be any reason for it. Either the DRAC will see the KVM and not be able to administer it (like what is happening right now), or the DRAC won't see the KVM at all.
It is very frustrating, and Dell's techs seem apologetic, but there doesn't appear to be a fix for it. It just sucks.
Next time I'm headed into the colocation, I'm just hooking the console port into the KVM that is connected to the non-blade servers. At least that way I'm not reliant on Dell's sorry excuse for a controller to access the video on my servers.
A) Dell's DRAC is a very useful device which facilitates remote administration
B) at least it would, if it didn't suck so much
The blade enclosure comes with a DRAC module that is inserted into a slot in the back. It's paired with an avocent KVM module connected to the blade units. You access the KVM through the DRAC web site, which is the real problem.
Every Dell technician I've complained to has said the same thing. "Yes, the DRAC is slow. Very slow, and underpowered". It's not just slow, it vacillates between borderline and completely unusable. On a good day, expect 3 minutes for the page to load. On a bad day, don't expect to load the whole page.

It also seems to occasionally lose track of the KVM. It's happened a few times so far, and there doesn't seem to be any reason for it. Either the DRAC will see the KVM and not be able to administer it (like what is happening right now), or the DRAC won't see the KVM at all.
Next time I'm headed into the colocation, I'm just hooking the console port into the KVM that is connected to the non-blade servers. At least that way I'm not reliant on Dell's sorry excuse for a controller to access the video on my servers.
Wednesday, October 29, 2008
A couple of amusing videos...
This is hilarious...I can't believe I only just found this
The Website Is Down
and..I...I'm not even sure what to say to this....
why you should give your sysadmin a day off
The Website Is Down
and..I...I'm not even sure what to say to this....
why you should give your sysadmin a day off
Tuesday, October 28, 2008
Outsourcing your web hosting
I found this link on Reddit today, and I thought that some of you may be shopping around for hosting. 4 things your web host doesn't want you to know
At one point in time, we were looking at a hosted space away from our (then pitiful) primary and backup sites, the idea being that as a last case scenario, we could redirect users to that site which stated that we were having issues.
Eventually we got to the point that we felt comfortable with our data sites and canceled the hosting, but if you don't currently have a backup solution, a cheap parked host somewhere might not be a terrible idea.
Here's a page on how to choose a web host. There are many similar pages
At one point in time, we were looking at a hosted space away from our (then pitiful) primary and backup sites, the idea being that as a last case scenario, we could redirect users to that site which stated that we were having issues.
Eventually we got to the point that we felt comfortable with our data sites and canceled the hosting, but if you don't currently have a backup solution, a cheap parked host somewhere might not be a terrible idea.
Here's a page on how to choose a web host. There are many similar pages
Monday, October 27, 2008
The weather outside is frightful...
but the warm air blowing out of the back of the servers is so delightful...
It snowed on my way into work today. So of course, during my drive, I thought of all the things I was going to have to start doing. Among them was annual maintenance on the work generator.
Typically, maintenance on a largish generator is done on a per-hours-run basis. My manual says to check the oil every 8 hours of running, change the oil every 100 hours, and change the spark plugs every 500 hours. There are other things that should be done annually, however, and for those, we hire a local contractor to come over and take care of things.
In central Ohio, we can have nasty winters, but they don't typically start early with the bad storms. Since the daylight savings time has migrated to the first Sunday of November (this year, Nov 2nd), it makes a great reminder for me to schedule the service call. People in different climes may have to use different milestones, but whatever you use, use something.
I wrote in June about a generator failure that I don't want to experience again, so learn from my mistake and perform regular maintenance on your equipment.
If you haven't done it yet, now is as good a time as any!
It snowed on my way into work today. So of course, during my drive, I thought of all the things I was going to have to start doing. Among them was annual maintenance on the work generator.
Typically, maintenance on a largish generator is done on a per-hours-run basis. My manual says to check the oil every 8 hours of running, change the oil every 100 hours, and change the spark plugs every 500 hours. There are other things that should be done annually, however, and for those, we hire a local contractor to come over and take care of things.
In central Ohio, we can have nasty winters, but they don't typically start early with the bad storms. Since the daylight savings time has migrated to the first Sunday of November (this year, Nov 2nd), it makes a great reminder for me to schedule the service call. People in different climes may have to use different milestones, but whatever you use, use something.
I wrote in June about a generator failure that I don't want to experience again, so learn from my mistake and perform regular maintenance on your equipment.
If you haven't done it yet, now is as good a time as any!
Thursday, October 23, 2008
Multiple privilege levels in IOS
I knew it was possible to set up multiple privilege levels in IOS, I just never had the gumption to research how. It's not a topic that comes up a lot, just something that sort of sat in the back of my mind making me wonder how they did that.
If you're in the same boat as me, wonder no longer. A tutorial from ciscozine has you covered.
Just thought I'd throw this out there since I had always wondered
If you're in the same boat as me, wonder no longer. A tutorial from ciscozine has you covered.
Just thought I'd throw this out there since I had always wondered
Issue remote commands to Windows machines without installing ssh
The other day, I ran into a new (to me) blog that I'm going to start reading. It's called Boiling Linux and Windows (they also do some AIX). If you're a cross platform admin like I am, it's probably worth your while to check it out. I know I'm very light in Windows experience, so it's an educational site for me.
Anyway, last week they covered running commands on a remote host from a Windows machine to another Windows machine. The tool used is psexec, which seems like a combination of rcp and rsh. I say that because I don't actually think the communication is encrypted, according to this conversation. Still, it's an interesting idea.
Anyway, last week they covered running commands on a remote host from a Windows machine to another Windows machine. The tool used is psexec, which seems like a combination of rcp and rsh. I say that because I don't actually think the communication is encrypted, according to this conversation. Still, it's an interesting idea.
Wednesday, October 22, 2008
Outage Statistics
I started this blog to share information, and I love it when other people do the same thing. In our industry, it's often hard to come by statistical information regarding infrastructure issues, at least until they happen to us. To save us some time, Michael Janke has posted the last two years worth of network outage analysis over at his blog.
I won't post spoilers, but it's an interesting read that you should definitely check out.
I won't post spoilers, but it's an interesting read that you should definitely check out.
Tuesday, October 21, 2008
Does 'onboarding' sound like 'waterboarding' to anyone else?
The company I work for is small. Small enough that there's no HR department, and no standard procedure for bringing on new hires, or in the irritating parlance of the day, 'onboarding'. Given that we're about to bring several people in soon, something needs to be done to plan for it, so I'm going to take care of it.
I'm working on a standardized method for getting the person set up from the IT perspective, and I'm going to be getting input from the one HR-ish person who deals with everything like that. I think it'll make things a lot smoother, especially when it comes to ordering licenses and hardware, and setting up accounts and so forth. I don't think I'm to the point that everything can be scripted, but I'm getting a bit closer.
I also want to produce some documentation that educates the new employee about the internal processes and jargon. It's complex to the point that it takes most people a year to pick up the various interconnections between systems. I think I can make a massive cut in the time by providing diagrams and documentation. At the very least, it hasn't been tried before, so I think it would be an improvement.
Anyone want to share what kind of..."onboarding" *shudder* procedures they have at their business?
I'm working on a standardized method for getting the person set up from the IT perspective, and I'm going to be getting input from the one HR-ish person who deals with everything like that. I think it'll make things a lot smoother, especially when it comes to ordering licenses and hardware, and setting up accounts and so forth. I don't think I'm to the point that everything can be scripted, but I'm getting a bit closer.
I also want to produce some documentation that educates the new employee about the internal processes and jargon. It's complex to the point that it takes most people a year to pick up the various interconnections between systems. I think I can make a massive cut in the time by providing diagrams and documentation. At the very least, it hasn't been tried before, so I think it would be an improvement.
Anyone want to share what kind of..."onboarding" *shudder* procedures they have at their business?
Is RAID 5 a risk with higher drive capacities?
There's a very interesting discussion going on over at ZDNet about RAID5 and hard drive capacities. The premise of the discussion is that unrecoverable read errors are uncommon, but statistically, we're approaching disk sizes where it will start to matter. Here's a quote from the blog entry:
"SATA drives are commonly specified with an unrecoverable read error rate (URE) of 10^14. Which means that once every 100,000,000,000,000 bits, the disk will very politely tell you that, so sorry, but I really, truly can’t read that sector back to you. One hundred trillion bits is about 12 terabytes. Sound like a lot? Not in 2009."
That would mean a bad block when trying to read. It wouldn't be such a problem, except when it happens while you're rebuilding a RAID array after a drive failure. New drive failures are 3% for each of the first three years, after that, the rates rise quickly, according to that author and Google, who he referenced for the numbers.
So the problem becomes a RAID 5 array with a drive failure. Pull the disk out, add a new one in, and the array has to rebuild. Once every 12TB on average, that rebuild will fail, according to statistics.
Commenters have pointed out that the loss of a single block doesn't necessarily mean the array can't rebuild, just that the non-redundancy means loss of that particular bit of data. With backups, you can restore the individual file and have a functioning array. I think it would depend on the controller, but I don't have any data to back that up.
The author argues in favor of more redundant RAID mechanisms. RAID 6 can tolerate the loss of two drives, and other raids can lose even more, depending on the particular failure.
Just the other day, I had a RAID 0 fail, but that was from the controller dying. Have you ever had an array die during rebuild? How traumatic was it, and did you have a backup available to recover?
Also, if you could use a RAID refresher, I mentioned them a while back.
"SATA drives are commonly specified with an unrecoverable read error rate (URE) of 10^14. Which means that once every 100,000,000,000,000 bits, the disk will very politely tell you that, so sorry, but I really, truly can’t read that sector back to you. One hundred trillion bits is about 12 terabytes. Sound like a lot? Not in 2009."
That would mean a bad block when trying to read. It wouldn't be such a problem, except when it happens while you're rebuilding a RAID array after a drive failure. New drive failures are 3% for each of the first three years, after that, the rates rise quickly, according to that author and Google, who he referenced for the numbers.
So the problem becomes a RAID 5 array with a drive failure. Pull the disk out, add a new one in, and the array has to rebuild. Once every 12TB on average, that rebuild will fail, according to statistics.
Commenters have pointed out that the loss of a single block doesn't necessarily mean the array can't rebuild, just that the non-redundancy means loss of that particular bit of data. With backups, you can restore the individual file and have a functioning array. I think it would depend on the controller, but I don't have any data to back that up.
The author argues in favor of more redundant RAID mechanisms. RAID 6 can tolerate the loss of two drives, and other raids can lose even more, depending on the particular failure.
Just the other day, I had a RAID 0 fail, but that was from the controller dying. Have you ever had an array die during rebuild? How traumatic was it, and did you have a backup available to recover?
Also, if you could use a RAID refresher, I mentioned them a while back.
Subscribe to:
Posts (Atom)


