
AI, AI everywhere- or so it seems at the moment. Our insatiable appetite for the new capabilities and possibilities it delivers, and the seemingly continual arms race between the major models, has led to an explosion in data centers and a demand for hardware and electricity to train and serve these models.
What this means for people like you and me and for organisations is that ‘core’ hardware components have got expensive; the graph below shows the price of RAM (specifically DDR5–6000 2 x 32GB) over the last 18 months

And similarly, the price of NAND storage (specifically M.2 NVME 4TB):

I’m not even going to show what it has done to GPUs, as someone still suffering from the sticker shock of my RTX 3070 when I purchased it back in 2020 (a card I am not in any rush to replace!).
The reality at the moment is that organisations such as the big hyperscalers and AI shops are using their substantial means to purchase hardware at huge bulk directly from manufacturers, often bypassing OEMs like Dell/EMC, HP and others in the process. Many of them, including Google, are even working directly with chip fabs on manufacturing custom designs, but more on this point later.
Another key factor is the price of electricity, which in many Western countries is really high and only getting worse as data centers are popping up everywhere.
What this has meant is that many organisations running their existing data center footprints are feeling the pinch. A couple of years ago, I was working with a large UK retailer for whom their data center electricity bill was the same cost as hosting all the infrastructure in GCP; that’s before you even consider hardware refreshes and ongoing maintenance, important factors when you are considering TCO.
Why organisations should consider Google Cloud rather than continuing on-prem
Firstly, I appreciate my strong bias as a Google Cloud Architect and Engineer, and whilst my bias is strongly Google, some of these points are also relevant for other hyperscalers. I will look to answer each of the challenges I have raised above.
The hardware purchasing problem

Now the cynical side of me wonders whether the hyperscalers and AI shops have, in part, created this problem, much like how De Beers has controlled the diamond market. However, the reality is that supply-demand economics, limited manufacturing capacity (without the ability to quickly scale up), and relative price inelasticity have all combined to create the environment we are now in. The hyperscalers do have the distinct advantage in economies of scale, which makes them more attractive to work with than end consumers (as we have, for example, seen with Micron closing its Crucial brand).
What I also think makes hyperscalers interesting is their ability to work directly with the likes of chip fabs and even develop custom designs. Google, for example have custom designed hardware right through their stack, with the most prominent examples including the TPU (Tensor Processing Unit), Axion (based on Arm Neoverse V2), and TOPs (Titanium Offload Processors). This list is far from exhaustive, though.

As someone who has started choosing Axion as my default choice for my new (compatible) workloads (spurred in part by this excellent article from Dmitri at loveholidays), the option to use arm64 over x86 and gradually move workloads from the latter to the former without substantial capex costs is appealing. I think in a high-energy-price market, performance per watt has become a critical measure, which brings me nicely to my next point.
Electricity Costs
Hyperscalers and Google especially have spent considerable effort and resources optimising their data center footprints to reduce electricity consumption; this has resulted in some of the lowest PUE (power usage effectiveness) figures in the industry:

This low PUE figure translates into real savings during the operational life of hardware, ensuring that as little electricity as possible is used. This also, when combined with renewable energy, translates into low CO2 emissions. Though for most organisations I suspect the main draw will be the effect it has on the balance sheet.
Predicting the future

Whilst none of us can accurately predict the future, especially in the world of technology. It is clear, at least in the short to medium term, that AI will continue to disrupt. Whilst Moore's Law may no longer be holding true in the way it once was, I think customised ASIC chips like TPUs will continue to push the capabilities of computing; developing this type of custom silicon at small scale is likely to be prohibitively expensive for many organisations.
The other main challenge I see with predicting the future is predicting future workload requirements; this has always been a problem with capex-oriented models where equipment is often bought on 3–5 year cycles, with hefty ‘lock-in penalties' should you need to go back to the equipment OEM mid-cycle for additional capacity. Whilst these cycles are often being extended in the name of ‘sweating assets’, this can sometimes hamper an organisation's ability to deploy cutting-edge technology.
Generating TCO Reports for On-Premise Infrastructure
Fortunately, Google has made it really easy to create TCO reports, and with the latest updates to their Migration Center platform, even use AI to help with rightsizing, as per this demonstration from Cloud Summit London:
https://medium.com/media/63f435841c2dc2173df3925042db1ea8/href
On the topic of AI, I have also had really fruitful experiences with the ‘App Modernisation Assessment’ (formerly called codmod), which uses Gemini to generate really helpful reports on how to potentially modernise your existing applications to Google Cloud by analysing your source code.
Conclusion
In all, I don’t think there has been an easier or a better time to consider migrating your workloads to Google Cloud. The tooling to help you get there is only getting better at helping you understand your existing workloads, and with Google being at the forefront of AI development, the opportunity to combine existing systems with the latest AI innovations can help organisations gain a competitive edge.
Has the AI boom changed data center TCO and made cloud migration more attractive? was originally published in Google Cloud – Community on Medium, where people are continuing the conversation by highlighting and responding to this story.
Source Credit: https://medium.com/google-cloud/has-the-ai-boom-changed-data-center-tco-and-made-cloud-migration-more-attractive-641ec8ff70eb?source=rss—-e52cf94d98af—4
