Back from Hiatus – Summary Update 1

October 06, 2009 | Tony Bain

Here is a summary of the key discussions I have had over the last month.  Keep in mind, I’m no analyst.  This is largely opinion based on various conversations I have had with the relevant companies (for analyst insight see Curt Monash).

KickFire

I think Kickfire has been doing it a little tough lately.  The difficulties in a startup launching a hardware appliance (and associated logistics) combined with being too focused on the MySQL customer base has impacted the growth of this interesting start up.  But they aren’t taking it lying down and have adjusted the strategy and have added a new appliance to the range.  Kickfire now seems to have a stronger focus on the enterprise and has released a larger version of its appliance to provide a growth path.  As I have said all along, the MySQL aspect of their product is interesting but the solution as a whole is much more interesting and has much broader appeal than just the current MySQL customer base.

Flipping hardware appliances is a much tougher play than software only solutions, partly due to it being much more difficult for customers to get their hands on your stuff and have a play before they buy.  Hopefully Kickfire has mitigated most of these issues now though their online, on demand evaluation host.  I haven’t yet played with this but it is on my list of things to do over the coming month.

Kickfire’s enterprise strategy is just one of many that will be re-enforced by an Oracle acquisition of Sun.

Greenplum

Greenplum has addressed a perceived chink in its amour with the release of its column store capability.  Greenplum has taken the popular hybrid approach which means on a case by case basis you can decide if a particular table should be row or column orientated.  But as Daniel points out, it is a storage level only solution.  The storage only approach brings just part of the benefit of columnar stores, to achieve the full benefit the query execution engine needs to be aware of this layout (so features such as lightweight compression can be effectively used).  But I am sure this is an area where Greenplum will make further improvements in the future.

Groovy

Groovy has been working hard carving out its niche in the real time web data market.  If you don’t recall, Groovy makes an in-memory RDBMS that has been extended to provide real time data streaming capabilities.  Groovy has been positioning this into the large web properties who are working on creating new large scale, real time applications for their user base.

Aster Data

Aster has put out a number of announcements over the last month and I am trying to keep up.  Firstly they announced their tight integration with Hadoop.  This integration with Hadoop is map-reduce on the outside of the Aster Data platform (which apparently they didn’t have already although I think everyone assumed they did given their strong in database map-reduce message).  Aster has been banging the map-reduce drum for some time and is clearly the point of difference they are focusing on. 

Aster has also release version 4.0 of their platform a couple of days ago, then a few days ago I was a bit surprised to see an email from them referring to their platform as “the World's First Massively Parallel Data-Application Server”.  This seems to be a new name reference to the in database map-reduce stuff, maybe as an effort to differentiate themselves from the myriad of competitors in this space they are trying to carve out a new category all for themselves.  For me, the external map-reduce stuff makes sense as I can see this
being useful for data preparation on the way in to Aster and data
dissemination of data on its way out of Aster.  But I still don’t have
in my head clear examples when their in database map-reduce stuff is
useful.  I am sure it is but I have a feeling it is valuable on a case
by case basis which is difficult to articulate especially as a point of
difference message.  But I missed Curt’s map-reduce webinar (at the
last minute) so maybe that would have shed some light.  Anyway, they are running a webinar on this which you can register for here.

To me, Aster is more aggressively driving their platform into green fields trying to leverage their technology to find new customers and new markets.  Greenplum on the other hand is more ‘steady as she goes’, focusing on a more traditional and conservative enterprise data warehousing market (while still innovating ahead of the general purpose behemoth's).  The risks are on both sides.  When trying to define a new market you risk not finding one or finding one that is too small or “niche” to support your business.  With the conservative approach you risk being lumped in with everyone else, and in data warehousing ‘everyone else’ is now quite a long list. 

Is the RDBMS doomed (yada yada yada) ?

September 22, 2009 | Tony Bain

Ladybower PlugholeImage by Snooch2TheNooch via Flickr

I was speaking with Michael Stonebraker this morning.  I mentioned that lately many have been referencing comments he has made over the last couple of years.  And I also mentioned that many had interpreted them as he was implying the RDBMS is “doomed”.  Mike has been saying the same thing for years, but the current NoSQL movement seems to have picked up on this and highlighting one of the RDBMS's own pioneers is predicting its demise.

I asked Mike to clarify this.  My interpretation of his response is as follows.  I understand that he doesn’t believe the relational database itself is doomed.  Instead the current general purpose implementations, or “elephants” using his words, were out of date.  By moving away from a historical GP function into something more specific in focus, either in transaction processing or analytics, you can easily get 50x performance improvement over GP RDBMS.  This doesn’t necessarily mean moving away from the “relational” nature, but instead changing some core design principles for how a RDBMS is implemented.  It is this improvement factor that will see “new” specialist platforms overtake “old” general purpose platforms.  That is gradually, over time.  However Mike also mentioned the relational data model doesn’t make sense in a number of disciplines, particularly in sciences, and alternative modeling paradigms will offer benefits to this market (hence his focus on SciDB).  So while relational is a valid data model, other data models are also needed.

I have a similar position to Mike, but perhaps with a few differences. 

– Firstly I agree with the mantra that current GP RDBMS platforms provide only a “middle of the road” capability, and we gone too far in using a GP RDBMS for everything.  However I do believe there is a long term future for the GP RDBMS.  A general purpose application requirement will continued to be well suited for a general purpose platform.  With a specialist only approach, a general purpose requirement may need both a specialist OLTP platform and a specialist Analytics platform to provide the same capability.

– I agree that with an extreme requirement, either analytics or transaction processing, a specialist platform is well suited.  But I don’t see the choices of just MPP or memory resident RBDMS as being a broad enough set.  Apps that use a db just as a persistence cache will benefit from a high performing, scalable database platform with much tighter integration with the object model.  I am not sure any of the current NoSQL platforms have it quite right yet, but when these guys eventually get together with the database guys and work on these things together they may get there.

– I don’t think a 50x performance speed up on its own is enough to drive change in OLTP.  I have written before how difficult it is to get into this market and how tight Oracle, Microsoft & IBM have this sewn up.  But I don’t believe it is impossible, I think you just need to bring slam dunks on multiple fronts (performance just being one of them).

Anyway I feel like I am a bit of a broken record at the moment.  I have been addressing the “is the RDBMS doomed” question a couple of times a day for some time. Time to focus on something else for a bit.

Some Initial Thoughts on Oracle Exadata V2

September 16, 2009 | Tony Bain

There will be plenty of detailed coverage on Exadata V2 so I won’t attempt to replicate that.  However I do have a couple of initial thoughts which I would like to share.  For those who missed it, Oracle has just announced Exadata V2 (which is their pre-built “machine”).  Exadata V1 was built using HP equipment, Exadata V2 is using Sun.  The main addition to Exadata V2 seems to be an extra tier in the memory hierarchy, a flash cache.  Oracle is very quick to point out this is not flash disk, but it is flash memory, Sun’s FlashFire technology (flash disk or SSD’s was always going to be a transition technology, flash memory doesn’t have the physical constraints of moving parts disk so the whole “disk” concept for flash doesn’t make too much sense other than it fits easily with current architectures).

The new memory layer (Processor Cache’s -> DRAM -> Flash Cache -> Disk) coupled with Oracle’s algorithms to effectively use the Flash Cache layer brings performance benefit to the solution (+ all the other improvements 12 months of hardware innovation brings, faster CPU’s, more memory etc).

My initial thoughts are:

  • Kudos to Oracle.  They are the first vendor to really bring a bunch of this leading edge technology together in a semi-mainstream way.  Flash Cache, Inifiband interconnects, DBMS optimizations using flash hasn't really surfaced anywhere outside of startups yet.
  • So what happens to Exadata V1 customers using the HP solution?  This is only about a year old.  Some analysts are suggesting there has only been minor sales of Exadata V1 (I am not an analyst so don’t really know).  So why would HP continue to support a platform where no new sales will be created, when potentially only a limited number of customers have it today?  Possibly Oracle will offer attractive terms to move existing HP Exadata V1 customers to Sun Exadata V2.
  • It is a preconfigured solution that you by in certain size configurations.  Small, half rack, full rack, multiple racks.  I think Larry said that 3 racks will give you a PetaByte of storage capacity.  This is fine, except they are targeting it for use with OLTP and data warehousing workloads.  It seems odd that to get very high computational resources for transaction processing, you would also get massive volumes of potentially unnecessary storage capacity.  It will be interesting to see if they allow the balance between processing & storage to be modified as part of configuration.

I have had some questions along the lines of “isn’t this back to the one size fits all approach?”  Well yes it is, but Oracle never really moved away from this in terms of the core DBMS.  It is my understanding that Oracle Exadata was still the general purpose Oracle DBMS & RAC but on a hardware platform optimized for accessing large data sets (making it a data warehousing solution).  Using FlashFire, the hardware can now do high levels of random I/O (I think 1m random I/O’s was quoted) which makes the hardware platform general purpose as well.

One interesting question will be if, under Oracle, other vendors can buy the exact same hardware configuration from Sun and optimize their DBMS for Flash also?   If so, it may be difficult for them to do this in a way that is price competitive.  And will competitive DBMS vendors really want to help fill Oracle’s pockets further?

If we expect to see more of this hardware alignment between DBMS vendors where does that leave Microsoft?  Maybe HP is already peeling the Exadata V1 logos off their racks and sticking Microsoft Madison logo’s in their place?

UPDATE:

Oracle has put out a FAQ which partly answers some of the questions.

OLTP back into focus

September 14, 2009 | Tony Bain

I haven’t blogged in over a month now.  This is for a number of reasons.  Firstly I have been flat out with various activities.  This included a trip to VLDB in Lyon mid month.  Secondly, a lot of the companies I have spoken with this month aren’t ready to speak publically so hence no blog posts resulting from these sorts of discussions.

However there has been a wiff of a change in the air in terms of focus that is interesting and worth highlighting.  After years of lots of innovation around data analytics, OLTP is starting to make a comeback in terms of reclaiming some of the limelight.  Much more on this between now and the end of the year, but a couple things to watch:

VectorWise

August 01, 2009 | Tony Bain

I was fortunate enough to speak with Marcin Zukowski earlier about VectorWise.  If you missed it, VectorWise came out of stealth mode a day or two ago.  The have announced a joint partnership with Ingres and essentially are claiming impressive analytic RDBMS performance gains on conventional hardware.

To start with, a key message that I think needs to be communicated here is that this is not a product announcement.  Ingres and VectorWise have announced a partnership in which they of course plan to build products together, today those products are still in the works.

VectorWise is a spin out of CWI based on research that was undertaken by Marcin and others, research that centered on MonetDB.  Explaining the essence of VectorWise is difficult because it is largely internal DBMS data storage & processing logic, but I will have a go.

The modern RDBMS is based around design principles that stem from general purpose OLTP roots and historical hardware architectures (this is partially true even for some of the newest analytic platforms).  These design principles in a nutshell focus on the fact that disk is slow & CPU is fast.  Data is seeked or partially scanned off disk and cached.  Row-by-row (tuple-by-tuple) operators process that data, passing the outcome of each operator to the next as part of a queries execution plan until ultimately producing the result. 

Traditionally I/O is the main bottleneck, so to make the database faster you add more I/O bandwidth.   Today, disk requirements may be up to 100x the actual capacity needs, so many disks are necessary to achieve the I/O bandwidth to provide performance for an analytical RDBMS implementation.  Even though the RBDMS’s may parallelize query operators across cores, this typically works by partitioning data between cores, yet each is still processing on a tuple-by-tuple basis.

Conventional wisdom?  Well maybe.  You see disk is only really “slow” when it is doing random seeks.  Give a disk something sequential to do on the other hand and things are very different.  Modern disks are able to sequentially scan in the range of 150MB per second.  An array of 10 disks should therefore be able to return sequentially read data in the range of 1GB per second. 

When it comes to databases, column based storage has been found to effectively structure data for a) high levels of compression and b) sequential access.  VectorWise makes use of both of these technologies to help it achieve high levels of sequential I/O.  The problem now however is that disk may no longer the bottleneck.  While we can get 1GB a second sequentially off disk relatively easily & cheaply, processing tuple-by-tuple at this rate is very difficult.  As it turns out, a RDBMS’s may only achieve a data processing rate of 50MB a second per CPU core.  This makes the CPU processing limitations a big bottleneck for analytics data sets, assuming the above figures we would need over 20 cores to keep up with 10 disks (and of course CPU cores don’t scalability linearly).

If we step out of the database world for the moment into the world of high end computer games, or high end scientific processing, we find their use of current CPU technology is much more advanced than what we are used to.  They are using new CPU extensions (MMX, SSE, SS2, SSE4.2 etc) to parallize & pipeline computation within a CPU’s core meaning they are processing orders of magnitude more instructions per core that what a traditional RDBMS typically has been able to. The exact details are too low level to discuss here (many of the research papers are available online) but it is fair to say, modern CPU architectures contain advanced features that to date haven’t effectively been exploited by database vendors.

Enter VectorWise.  Their aim is to marry storage technologies which allow high levels of sequential I/O to occur with query processing logic which is designed for modern CPU architectures.  Rather than process tuple-by-tuple they are processing “vectors”, groups of tuples, leveraging modern CPU extensions and high levels of on-chip cache to allow the CPU to carry out higher data processing throughput.  The result is instead of the 50MB a second in a tuple-by-tuple approach, VectorWise are able to achieve processing rates in the range of 500Mb-1GB a second per core in some situations.  This means processing rates of 8GB a second or more could be possible with relatively low end hardware.

“In some situations” is the key point to stress here, this obviously isn’t a blanket gain that applies to all analytic data sets, workloads and query requirements.  Just what those situations are will be the key to their technologies success, how well it actually applies to real world data sets and queries.  I wouldn’t expect to see too many specific examples on this until a product beta appears.  But the theory is VectorWise can offer high levels of processing capabilities with existing mainstream hardware.  At this point VectorWise isn’t even focusing on MPP instead they are single node focused.  If their scalability claims pan out you can imagine how this could allow a single node solution to be competitive with existing low to mid scale MPP solutions that are based on a more conventional query processing architecture.

This isn’t VectorWise’s only trick up their sleeve.  They are also are leveraging research around column based storage, compression, piggy-backed (shared) scans and so on.  Much of the research that has been adopted by VectorWise is referenced from their web site.

So VectorWise have impressive technology, so why then partner with Ingres rather than a larger vendor (or going at it alone)?  Marcin offers a few reasons.  Firstly, as academics they feel strongly that open source is cool so this path was greatly preferred over a relationship with a non-open vendor.  Secondly Ingres will allow them to deliver their technology in an uncompromised fashion.  Marcin mentioned that if they had partnered with one of the big three vendors, that vendors existing product strategies and investments would have likely meant their ideas could have only been implemented in partial form.  Ingres on the other hand is going to allow them more of a green field.  And of course, a partnership with Ingres makes sense from a go to market perspective as Ingres already has a worldwide reputation, a global customer base, sales & marketing capabilities etc.

Marcin confirmed that Ingres have an exclusive license to their technology, and first option to acquire them for a certain period of time.  This allows Ingres to really invest in the relationship without the fear of the carpet being pulled out from under them. 

VectorWise clearly are applying innovative research to analytical RBDMS requirements.  But as interesting as the technology sounds, the proof in the pudding will be how well these design principals translate to real-world analytical processing requirements in mainstream product form.  This remains to be seen, but Ingres and their community clearly has high hopes.

VectorWise is clearly differentiated when comparison with a traditional mainstream RDBMS running on mainstream hardware.  However in this current market we have lots of different approaches to the problems described.  Kickfire for example use their own SQL Chip processor to increase data processing rates and other appliance vendors are using FPGAs etc for similar purposes.  The comparison of these different approaches and the relative effectiveness of each approach still need to be examined, however a mainstream hardware approach has obvious benefits.

Maria Update

July 31, 2009 | Tony Bain

MySQL PanelImage by Sebastian Bergmann via Flickr

I had a quick chat with Michael Widenius today.  He is on vacation so tried to keep the call short.  Essentially spoke about two topics, Oracle & an update on Maria.

The Monty Program has 15 staff now.  Their focus is getting the MariaDB branch of MySQL ready for release, I understand they have a target of next month (August) for this release.  The Maria storage engine has been delayed for the time being with the focus being on the branch release instead.  PBXT and XtraDB will two of the storage engines included in this release.

The Open Database Alliance is a key initiative on the Monty Program, started in conjunction with Percona.  Essentially the ODA is a network of third part MySQL services organizations with an operating agreement between them.  The idea is to build a credible global support capability for MySQL outside of Oracle/Sun.  But also to provide the same structure to other open source database initiatives, such as Postgres.  If you are in the business of providing MySQL servies, or services to other open source database platforms, I suggest you check it out.

I think we shouldn’t have to wait too much longer before Oracle is in a position where it can start talking about its MySQL plans.  At this point however I would think that MySQL's path forward could be quite different to its recent past.  I originally planned to expand my thoughts on this here, but as this is and can only be speculation (as Oracle hasn’t made an official comments yet) maybe it is just better to wait and see.

The NoSQL community needs to engage the DBA’s

July 30, 2009 | Tony Bain

The NoSQL movement has been gaining some steam lately, with discussion forums and mailing lists popping up all around the web.  Despite having a career that has been centered on the RDBMS, I have made no secret that I think we have gone too far down with our RDBMS for everything mindset.  I think we need to add a few more tools back into our data toolbox. 

Today, 99.5% of new data centric developments started will use a RDBMS by default.  Maybe .5 of a % will consider using something as obtuse as a NoSQL platform.  By experience I know the majority of people discussing NoSQL platforms today are web developers.  In fact there is almost a sense of trying to trying to keep this under the radar of DBAs.  If we don’t talk to the DBAs about this stuff then they won’t bother us with all that jabber about consistency, data integrity, robustness and recovery. 

Actually, many of the NoSQL projects are touting one of the key benefits of a NoSQL platform is you can do big data without the need of a costly DBA.

Baloney.

This shows me that the people making those comments have no idea what DBAs do and what happens with critical data applications post deployment.

A NoSQL data platform may have a different approach to operational management than a RDBMS, but a large part of the requirement will be the same.  It doesn’t matter if you have 10, 100 or 1000GB of data deployed on a NoSQL platform or an RDBMS.  Someone still needs to be thinking about backups & recovery, availability, capacity planning, performance monitoring, import/export, data integration, tuning & optimization, replication latency and so on.  Also, I have never come across any technology that works perfectly 100% of the time, so when things don’t work as expected and nodes are out of sync or partial data corruption occurs at 2am, someone will still need to fix it.  Guess who that is going to be.

DBAs are critical to any wide scale success with NoSQL platforms.  They need to be engaged and educated.  Sure they are going to be really annoying for quite a while, ripping into common NoSQL limitations such as lack of transaction support, eventual consistency, data duplication & application controlled data integrity.  However over time they will start to see the positive aspects as well and learn sometimes a mallet isn’t the only tool required.

HamsterDB

July 30, 2009 | Tony Bain

With all the noise over key/value stores recently, we should keep in mind that this technology isn’t exactly new.  It is being applied to new problems, but many of the foundations have been around for decades.  Probably the oldest of them all, Berkley DB came into existence during the mid ‘80’s and now has over 200 million deployments (according to the Oracle web site).

HamsterDB, while not having the same pedigree of Berkley, has been steadily worked on by Christoph Rupp for the last 5 years.  I spoke to Christoph yesterday about his release of a new edition of Hamster.

Hamster primarily has been a single threaded data store for embedded use, however Christoph has expanded Hamster into two editions, the second being a multi-user transactional platform.  This month sees the release of the BETA build of this edition.  Key Features of the NoSQL HamsterDB Transactional Key/Value Store:

  • ACID Compliance
  • Lock Free Architecture (transactions fail on conflict rather than block)
  • Transaction logging & fail recovery (redo logs)
  • In Memory support – can be used as a non-persisted cache
  • B+ Trees – supported but additional indexes are user maintained (see below)

Performance details are still sparse, but the embedded edition has done very well when compared with BerkleyDB.  HamsterDB is also licensed using a mixture of GPL and commercial licenses.  It provides native C++ support but wrappers exist for .NET, Java & Python.

Hamster is a pure key/value store so doesn’t have the feature set of the hybrids, such as MongoDB or Wistla, i.e. lookups are by key only, B-Tree indexes are supported but must be manually maintained by the application etc.  HamsterDB also doesn’t have a distribution layer so is for single node use.

But, from the perspective of a simple, lightweight, high performance key/value alternative HamsterDB looks very interesting.

HadoopDB discussion with Daniel Abadi

July 23, 2009 | Tony Bain



I spoke to Daniel Abadi this morning about his HadoopDB announcement that came out a couple of days back.  I am sure this has been a busy time for Daniel and his team over in Yale as HadoopDB has been getting a lot of interest which I am sure will continue to build.

Some notes from our discussion:

  • HadoopDB is primarily focused on high scalability and the required availability at scale.  Daniel questions current MPP’s ability to truly scale past 100 nodes whereas Hadoop has real examples on 3000+ nodes.
  • HadoopDB like many MPP analytical database platforms uses shared nothing relational database as processing units. HadoopDB uses Postgres.  Unlike other MPP databases, HadoopDB uses Hadoop as the distributed mechanism.
  • I am adlibbing here, but I understand that Daniel doesn’t dispute DeWitt & Stonebrakers (and his) paper which claims Map/Reduce underperforms when compared to current MPP DBMS.  HadoopDB however is focused on massive scale, hundreds or thousands of nodes.  Currently the largest MPP database we know of is 96 nodes.
  • Early benchmarking shows HadoopDB outperforms Hadoop but is slower than current MPP databases under normal circumstances.  However when simulating node failure mid query HadoopDB outperformed current MPP databases significantly.
  • The higher the scalability the higher the possibility of node failure mid query.  Very large Hadoop deployments may experience at least 1 node failure per query (job).
  • HadoopDB is usable today, but should not be considered an “out of the box” solution.  HadoopDB is an outcome from a database research initiative, not a commercial venture.  Anyone planning to use HapoopDB will require the appropriate systems & development skills to effectively deploy.

HadoopDB is an innovative approach to the scalability challenges that continue to push the architecture of the modern database forward.

Could MySQL be pigeon holed by Oracle love?

July 17, 2009 | Tony Bain

A while ago, about 16 years ago now, I had a desktop computer.  It wasn’t a PC.  It was an Acorn.  It had an ARM processor in it.  Despite the rest of the world starting going crazy for the new Pentium chip, the Acorn with its ARM processor could run rings about it in terms of computing power.  And it was simple and easy to use, I used to write applications in assembly code for it (and it didn't have a fan!).

Not too long after that Acorn went under, Arm was already off on its own to find a new market.  Its RISC technology was licensed in many different ways.  Despite some isolate cases where the technology was again used on the desktop or even in supercomputers, largely those licensing it didn’t require another desktop processor.  They needed a mobile processor, which ARM’s technology was great for too.  Over time the ARM processors have become well known for their mobile capabilities and their desktop & supercomputer capabilities became less widely known (or cared about).

So why am I telling you all this?

Well as we all know, Oracle has yet to make any public statements about their intentions for MySQL.  Sitting in the Hannah Montana movie with my kids (don’t ask) tonight I was thinking about possible scenarios that could play out.  One of the interesting ones is what happens if Oracle positions MySQL as an entry level database, or as small scale web backend database, and showers it with love and attention, sales & marketing effort in that space. 

Is it possible that MySQL could start to become known for that limited capability only and recognition elsewhere could start to fade?  Would it matter?  Would this make sense, how would this be advantageous to Oracle? 

Rhetorical questions really as I am just thinking out loud, just thinking out loud…