Showing posts with label improve. Show all posts
Showing posts with label improve. Show all posts

Sunday, 14 July 2013

How to monitor and improve the health of your datacentre

Your system administrators are looking at their screens, and all seems well. A sea of green denotes that every monitored system is performing optimally, and the system administrators’ thoughts are turning to the weekend and a 48-hour splurge of online gaming or code debugging.

On the helpdesk meanwhile, the screams of anguished users can be heard, bemoaning the fact that their productivity has plummeted – due to poor access to or poor performance of the IT platform.


Somehow, there has been a serious disconnect between the system administrators’ metrics and the users’ experience. From the IT department’s point of view, everything is running well – but are they looking at the wrong things, or perhaps not looking at enough of the right things? For the user, identifying blame is easy: it’s IT’s fault – but that is usually a gross, and often unjust, simplification.

Just how “healthy” is your datacentre? Is it doing what the business requires of it, or is it part of the problem? To answer this key question, datacentre managers must first identify what it is, exactly, that they need to measure.

To check the datacentre’s health and to minimise downtime, there is a need for multiple levels of measurement – from the highly granular, equipment-based monitoring and reporting, to the outside-in monitoring and reporting from the user’s viewpoint.

At the datacentre IT equipment level, monitoring its state and performance is no longer enough. Being reactive to problems is storing up issues – an “N+1” redundancy approach (having one more item of equipment than is truly needed) turns into an “N” approach if an item fails – and if a second one then fails while the first is being fixed, IT disaster strikes.

It is far more strategic to use a predictive approach – monitoring factors such as the temperature of key components such as the central processing units (CPUs) and disk drives; monitoring the power draw to see if this suddenly and unexpectedly alters or is trending upwards, and replace the system before it fails.

Datacentre managers should also understand that during a replacement, an N+1 approach is no longer providing any redundancy – so they should either go for an N+2 (or greater) approach, or ensure that the key component are easily accessible so that replacement can be carried out rapidly. This will help to minimise the time where redundancy is not in place.

Next is the environmental health of the datacentre. The use of monitoring tools for overall temperature, smoke and humidity, along with infrared heat sensors, will allow problems to be detected before they become issues. By linking these to the equipment monitoring systems, IT will be able to connect the presence of a datacentre hot spot (indicated by an infrared sensor) with a specific piece of equipment which can be swapped out or shut down to prevent the problem getting out of hand.

The broader facility and its equipment also need to be monitored and assessed to maintain a datacentre in good health. Whereas facilities management may be using a building information modelling (BIM) tool, this will generally not be integrated into IT’s systems management tools.

The use of a datacentre infrastructure management (DCIM) suite may pull everything together, but that alone will not suffice. In addition to implementing DCIM tools, the health of the facility’s power distribution, uninterruptible power supplies (UPSs), auxiliary generators and cooling systems have to be linked in to the overall view of how the datacentre is performing.

Through using modular systems throughout a datacentre infrastructure – from the IT equipment to the facility support equipment, failures of individual pieces of equipment can be allowed for.

Where possible, datacentre teams should use load balancing capabilities – for example, using intelligent virtualisation of servers, storage and networking equipment and intelligent workload management modes in UPSs and generators – to provide the maximum levels of business continuity.

Load balancing will provide much higher levels of availability than a direct, simple N+1 approach, as the failure of even two or more items can still be dealt with, even if the application performance is affected.

But assessing the datacentre infrastructure’s health is not just an IT question, it is a business question.

The above discussions deal with the datacentre itself – and for many organisations that may already have the above in place, it may be seen as being enough.

The problem is that the screens that the systems’ administrators are looking at are generally part of that datacentre environment, and are connected to the systems through datacentre networks at datacentre speeds. So it is hardly surprising that everything looks as if it is working well while the helpdesk goes into meltdown.

The datacentre is generally connected to the rest of the organisation through local and wide area networks (LANs and WANs). Users access the datacentre through these networks. If there are problems anywhere along these connections, the user will have a poor experience – and will contact the helpdesk, often with the perception that it is an application or a datacentre problem, rather than a network one.

It is incumbent on the datacentre manager, therefore, to be able to monitor connectivity across the different types of network to ensure that the datacentre serves the business users’ IT needs.

Any datacentre manager looking to fully support the business should take an end-to-end systems management approach, so ensuring that the organisation is working against a healthy IT platform – not just a healthy datacentre

Another challenge is that many of today’s workers are mobile and/or working remotely and will therefore be accessing the company’s datacentre services through public connectivity – ADSL lines or Wi-Fi or mobile wireless networks. Being able to measure network performance across these less predictable connections can be problematic – but providing tools to the helpdesk for them to “ping” the user’s device and see if latency, jitter or packet loss are causing issues will help the IT in effectively identifying the root cause of any issue.

The final area is around the user’s device itself. A PC may have a disk drive that is full, a tablet may have a process that has hung at 100% CPU utilisation, or a virus may be affecting overall performance. Putting in place tools that enable the endpoint device to be monitored and fixed – automatically wherever possible, or through efficient and effective means via the helpdesk where automation cannot be used – will again make root cause analysis easier.

The use of human mean opinion scoring (MOS) systems can also help. Rather than depend on technical measurement of the performance of systems where comparing a pseudo-transaction’s performance against an old service level agreement (SLA) and getting a green signal, asking real users as to their experience will be far more illuminating.

If users find performance to be too slow, it is no use pointing to the SLA and saying that it is within agreed limits – if the perception is that it is too slow, then it is down to IT to see if performance can be improved.

Monitoring the health of the datacentre is like monitoring the health of a person: focusing on just one area can mean missing where the real issue is and thereby lead to the failure to properly treat the problem.

Determining the set of measurements that will provide an accurate assessment of datacentre health in terms of business performance requires a holistic approach.

Any datacentre manager looking to fully support the business should take an end-to-end systems management approach, so ensuring that the organisation is working against a healthy IT platform – not just a healthy datacentre.

Clive Longbottom is service director at analyst Quocirca. The datacentre consultancy firm has three papers that cover ITLM and IT financing available for free download here: Using ICT financing for strategic gain; Don’t sweat assets, liberate them; and De-risking IT lifecycle management.

Friday, 12 July 2013

Gadgetwise Blog: A Sensor Steps Up to Improve Your Stride


Polar’s Stride Sensor sends signals to a Bluetooth Smart-equipped phone that monitors your progress while you run.
Polar’s new Stride Sensor allows runners to collect information about their stride and then link this information to a phone app that calculates speed and distance data through GPS. It connects to the phone through Bluetooth Smart, and is helpful for people trying to track calories or improve their running performance.

Using the device just a few times, I was able to take 30 seconds off my mile time by taking smaller steps. For a runner, that is pure gold.

The ovoid, $80 Stride Sensor is about the size of half a small hen’s egg, larger than the Garmin model, which connects through an Ant+ signal.

The Stride, which can be used in tandem with a heart rate monitor, sends signals to a Bluetooth Smart-equipped phone (later-model iPhones and Samsung Galaxys, according to the Bluetooth Web site). The phone, using the Polar training app, follows and charts your progress. While you are running, it gives audible guidance, telling you when you cross set distance markers, or when you reach a target distance or time, or burn a specific number of calories.

But before this can benefit the most serious runners, it needs some improvement. For maximum accuracy, the sensor needs to be calibrated on a one-mile course. No matter how many times I set it, it remained inaccurate by the same tenth of a mile. That is enough of an inaccuracy to make the sensor nearly useless to a competitive athlete. A company spokesman said Polar was looking into the problem, which may in fact be with the app, or the phone GPS.

Wednesday, 26 June 2013

Big data project aims to improve Dutch flood control, save the government millions

A big data project called Digital Delta aims to investigate how to transform flood control and the management of the entire Dutch water system and save up to 15 percent of the annual Dutch water management budget.

 

IBM will collaborate with Rijkswaterstaat, the part of the Dutch Ministry of Infrastructure and the Environment that is responsible for the design, construction, management and maintenance of the waterways and water systems in the Netherlands. The project also involves the University of Delft, local water authority Delfland and the Deltares Science Institute, the organizations said in a joint news release Tuesday.

 

They will investigate whether data gathered by more than 100 different projects in the Netherlands that deal with water management can be combined and be made accessible to accelerate water management innovation, said Djeevan Schiferli, an IBM Business Development Executive who is involved in the project.

 

If the research is successful, the system can be applied to other areas in the world, Schiferli said. The principle has already been discussed with governments in the U.S. in New York and New Orleans, as well as in Japan, South Korea and Australia, he said.

 

In the next 12 months, the Digital Delta group will investigate how to integrate and analyze water data from a wide range of existing sources, including precipitation measurements, water level and water quality monitors, levee sensors, radar data, model predictions, and current and historic maintenance data from sluices, pumping stations, and locks and dams, the group said.

 

Because 55 percent of the Dutch population lives in a location prone to flooding, every water-related event is critical and can impact businesses, agriculture and citizens’ daily lives, they said.

 

Due to this, the Dutch water management budget adds up to €7 billion (US$9.2 billion) each year, and costs are expected to increase €1 billion to €2 billion by 2020, unless something is done, according to the release. By combining data, Digital Delta thinks it can save up to 15 percent, Schiferli said.

 

Digital Delta wants to reduce costs by dealing with IT and the available data in a better way, he said. At the moment, parties use a third to half of their budgets to collect and disseminate the data, according to Schiferli. “And that is even before the data can be interpreted,” he added.

 

The initiative aims to provide water experts with a real-time intelligent dashboard to harness information so data can be shared immediately across organizations and agencies, according to the release. Involved organizations can use the dashboard to help prepare for imminent difficulties as well as enabling authorities to coordinate and manage response efforts. Over the longer term, the project aims to enhance the ongoing efficiency of overall water management, it added.

 

Digital Delta’s goal is to improve the management of the Dutch water system in many different ways. For instance, by modeling weather events, the Netherlands should be able to determine the best course of action including storing water, diverting it from low-lying areas, and avoiding saltwater intrusion into drinking water, sewage overflows and water contamination, they said.

 

One satellite company can, for instance, spot a 100 meter stretch of dike that is sagging, Schiferli said. If those responsible for tending to the dike know this, they can put sensors in just that stretch instead of in the whole dike that is, say, 10 kilometers long, he said. Monitoring a stretch of 100 meters is much cheaper than fitting the whole dike with sensors, he said.

 

While all of the relevant data to determine where to put sensors may now be available, it is not combined so not everyone has access to it and this leads to unnecessary expenditures, he said.

 

Another project in the initiative aims to link a sensor to a well in a traffic tunnel, Schiferli said. If the data of that sensor is linked with the maintenance system and weather data, an estimation can be made of how likely it is the well will overflow or clog when heavy rain is expected, he said. The system could then send a warning to relevant authorities to check the well, which may prevent a traffic jam in the tunnel, he said.

 

IBM’s Intelligent Water Software can be used for this. Rijkswaterstaat and local water authorities will manage water balance data and share the information centrally through the Digital Delta platform, the organizations said. This should make it possible for the Dutch water system to optimize the discharge of water and improve the containment of water during dry periods, and prevent damage to agriculture, the group said.

 

Also, a scalable early flood warning method will be developed by combining weather data and, water system simulation models and real-time measurement data from the water system, they said.

 

The Digital Delta project is aimed at water management. Similar systems could be developed to deal with droughts or for use in ship traffic, said Schiferli.

 

“The Australians, for instance, said they know a lot about forest fires—that knowledge could be useful to the Dutch when they have to deal with a fire in the dunes,” Schiferli said, adding that international cooperation is also a possibility.

 

The current research project costs €5.5 million and will go on for 12 months after which the governments decide if they want to go through with it, he said.