2014-02-18
The new processors and pop up ads
2014-02-11
The Data Warehouse's new Intel E5 V2 Processors
But what does these servers do?
2013-11-30
PHP pcntl threads, Intel E5 VS2 & check-MK
2013-09-08
Moving a Data Warehouse - 3
In some posts I have written about the migration of my Data Warehouse from my own hardware to Dell servers. This move was partly initiated by my transition to a new job in the company, the management didn’t dare to run the Data Warehouse on the hardware I built .
When I originally designed the Data Warehouse infrastructure two key principles were low cost and simplicity . I needed a database so I created a database server. I needed an ETL engine so I created an ETL server. Then I needed PhpMyAdmin so I created a PhpMyAdmin server. One function one physical server. The only extra in my Irons was an extra Network Interface for an internal server network. And in the beginning servers were scrapped IBM desktops. I maximized RAM, replaced the hard disk and added a network interface, dirt cheap. Then I installed a Linux and one application, fired up the new server and forgot about it for two to five years. My servers were mostly replaced when I needed more capacity not due to hardware failures. One of very few hardware problems I have had is described here .
I avoid software tweaking and optimization, I try do do standard installs right from the distro, e.g. I choose ‘big’ for Mysql config file that’s is about how much tuning I do. I have all databases and indexes on the same disk! About half a terabyte database with about eight million queries a day ( I have seen peaks over 15 million queries a day), this is on a custom made server with 16GB RAM.
Now this has changed with the migration to ‘real’ servers. The one server one function approach would have been all too expensive, so I had a choice either pack more functions in one server or go virtual. I have for some years wanted to test a virtual solution, so I decided to go virtual without testing. I decided I go for two servers one physical database server and virtual host for all other servers. I do not believe for a second you can have a low cost simple virtual high performance database server. But the rest of my Data Warehouse servers could well be virtual, this way I could keep my one function one server philosophy and still be reasonable cost efficient. I was right and I was wrong.
The new environment is much more complex, virtual servers add a software abstraction layer between the iron and Linux, and by going virtual you also need a software layer between your hard disks and the virtual servers for practical space management. For all this to work you need an expert to manage this environment. And an expert costs and the expert has his own preferences and experiences, e.g. a Linux professional does not necessarily know Mageia Linux. Since we didn’t have in house Linux operations expertise we hired a consultant. A mistake was not to listen to the consultants recommendation of virtualization software and Linux distro. Not that it’s difficult for a Linux professional to learn another distro it just takes some time, but more important, the support of my now non standard infrastructure it will always be exotic for the consultants operations team. I should have spent more time with the consultants upfront going thru the server setup. We would probably have had a better server setup still adapted to my Data Warehouse.
This is complication I would not have had if we had used Windows Server instead of Linux distros, since Windows is a singular opsys. I do not know if this is good or bad.
I end this post with a humble statement. Still few people seem to have my insights in hardware and infrastructure for Business Intelligence systems. Actually very few I talk to make any distinctions between any type of applications in this respect, the same hardware fits all give or take some RAM and CPU; that is the adaptation to applications you see. I believe most hardware infrastructure is grossly overpowered/priced and designed for ERP transactional applications.
2013-06-03
Moving a Data Warehouse - 2
Now we have moved the Data Warehouse to the new hardware. It is more than just a move to new servers, it is a complete redesign of the Data Warehouse. The mysql database server is still a physical server, but the rest of the servers have been virtualized . The data warehouse is now based on Ubuntu 12.04 LTS Linux, with one exception the heart of the Data Warehouse, the ETL server is a Mageia 2 Linux .
The Data Warehouse is in the middle of the blue circle (background photo Anders Nygård).
Here you see a schematic picture of the new Data Warehouse. The new Data Warehouse consists of three physical servers, the database server, the application server or virtual host and a satellite server in Japan. The database server connects to the DW applications via an internal switch and to external applications via some connections to the corporate network. The virtual host contains all virtual DW application servers. The ETL server has been divided into two server a communication server containing mail, ftp and a smb client, and the ETL server where all jobs inbound and outbound are run. The Japanese satellite server is a replica of the database. This new setup is much easier to maintain. I’m very grateful for the invaluable help Anders Nygård from Red Bridge provided. Without Anders skills and knowledge this migration project would not been finished for a very long time.
2013-05-16
Moving a Data Warehouse - 1
2013-05-03
New Data Warehouse Servers
I have written some posts about design and build your own BI infrastructure . Now when I’m getting a new assignment, we (my bosses :-) decided I had to replace my hardware with more professional irons.
The new Data Warehouse
I chose two servers from Dell, o ne PowerEdge R720 and one R720XD , both equipped with one Xeon ES-2609, 32GB 1600MHz RAM and 12TB RAID 6 SATA/Nearline SAS disk space ( I like oceans of space ).
I will use the R720 as a physical MySQL database server , and the R720XD will contain all other (virtual) servers. Going virtual is an architectural change I planned to do for a long time, but I never found the time. I rather would have used my custom build hardware, but the guys who decide decided otherwise. ‘Since you jump ship, we do not dare to use your old computers . We want proper servers.’ I do not know how the new infrastructure will perform but I got a hunch it will perform better than my now old irons. I will also go from 100MB network connections to gigabit, this will definitely speed up communications.
2012-07-16
Always online
2012-07-10
Business Intelligence and Hardware
Hardware is often neglected in applications design. In best cases you divide components of an application into separate servers, move the data to a separate SAN and add some extra RAM for performance and that’s it. The server infrastructure of applications is often well suited for transactional systems. Transactional systems do lot of random read/write of small chunks of data and very little processing on those small chunks.Business Intelligence systems on the other hand does not write very much, but reads a lot. Both reads and writes are mostly done in large chunks.
Since the middle of the nineties I have been interested in BI applications hardware infrastructure. My interest for hardware started with the spintronic revolution, I realized that hard disks (HDD) and memory RAM was going to be larger, faster and cheaper in future. HDD was the first in the spintronic wave. HDD had been too small, too slow and too expensive to allow for modern BI; with the SATA HDD we had cheap, fast and large HDD. The seek time (find the data to read or write) on these cheap SATA HDD is not impressive, so if you need to do lots of small random read/write it is not a wise choice. So the traditional servers still use expensive and small HDD with good seek time. These HDD are not very good for BI. BI are more interested in low transfer time (from disc into memory) than low seek time, and the SATA protocol can deliver more data than the discs can spin. More important SATA disks today are large, you can get 3 TB SATA disks for about 200€ and size matters for BI operations. SATA HDD are better suited for BI than more expensive and smaller HDD found in servers. Of course we will use Solid State Disks in a few years’ time, but still these disks are too expensive and too small for my liking .
Lots of RAM more than compensate for the slower SATA disks (compared with SAS and SSD), today you get 24GB SSD3 for about 300€. It is important to have enough of RAM to keep the active data set in memory to keep physical I/O low. Our BI database server needs 16GB to perform well. First day of the month I have noticed increased response times and some ‘peculiar’ MySQL behavior. I will install more memory and see if that helps. If not I have to find the root cause for the increased response times which most likely is missing or not good enough indexes. I still see recommendation about being careful with adding indexes. For BI systems this is wrong, wrong, wrong. You should sprinkle your database with indexes. All frequent queries should have optimized indexes. For the experienced DBA I recommend Relational Database Index Design and the Optimizers’ by Tapio Lahdenmaki and Mike Leach . This is serious, heavy and good reading about indexes. But I first try to mend performance problems with hardware, it is simpler to add RAM, than to analyze bottlenecks.
Processors today are so powerful any modern multicore processor will do just fine. I use high quality workstation motherboards. For my last database server I used an ASUS P9X79 DELUXE X79 S-2011 ATX motherboard for about 300€. It performs beautifully for our Business Intelligence system.
With such components it is an easy task to build high performance BI servers. I build the servers as simple as possible, this means I deliberately build them as single-point-of-failure. Simple server means few things that can crash and my servers are remarkably stable. The only things that have crashed so far are HDD raids and raid controllers (two times in eight years). Today disks are so big so we do not need to raid them anymore. Hardware development goes very fast; last year’s top notch hardware is ready for the scrap heap the next year. Using inexpensive servers give me the luxury to replace them more often. The normal lifetime for a server is three years and the life time cycle is test-production-backup-scrap heap. The database server I try to replace more often.
I’m aware of most experts do not approve of my ideas. But I created a working system after these ideas and it performs beautifully. I have put lots of efforts into my system; thinking, testing, measuring. More important than hardware is the database design. I have completely removed the traditional snowflake or whatever it is called design. The traditional BI database design patterns were conceived when hardware were expensive and disks were small. With today’s cheap hardware the old design patterns are a millstone around the BI system’s neck, we will see new simpler databases with more redundancies as in-memory computing becomes a reality. I should probably post about this, since I am an humble pioneer in this field.
This is how my ‘tin cans’ look 2012-12-23.