I got an opportunity to attend the NFD-32 event as a remote delegate. The two-day event was full of presentations on cutting-edge technologies related to networking.
Nile was one of the companies that impressed me with their innovative approach to technology and product offering.
Now ! if you do not know about Niles, it’s fine, as I didn’t know about them either.
However, it seems a disruptor in its own right.
It is a young startup that came out of stealth mode just last year-2022
And their unique product offering is Network-as-a-Service NaaS.
What is NaaS?
NaaS enables an enterprise to buy networking as a service rather than buying and maintaining its own networking hardware and software.
Does NaaS work for everyone?
Maybe not, and I bring this perspective to the blog’s conclusion. So stay tuned
What problem Nile is trying to solve?
In today’s rapidly evolving digital landscape, enterprises face numerous network challenges. These include the extensive time and money spent on choosing and purchasing hardware and software, requiring certified technicians for installation and configuration, and the time-consuming tasks of network monitoring and troubleshooting. All these challenges contribute to increased operational expenditures (OpEx) and enterprise complexity.
The solution is Network-as-a-Service or NaaS.
As a NaaS provider Nile removes this complexity by bringing its hardware, software, and management solutions and running the network for you.
What is included in Nile’s NaaS?
The following slide from the Nile nicely summarizes everything Nile offers.
The package includes everything from wireline switches to wireless access points, sensors to security. But infrastructure is not everything; they will plan, design, deploy, and maintain the enterprise network, thus covering the complete life cycle.
The benefits of NaaS are manifold. Firstly, it simplifies network operations by offloading Day 0 to Day N tasks, reducing the time and effort required for managing networks. This results in improved operational efficiency and reduced costs for enterprises. By leveraging NaaS, organizations can eliminate the complexities associated with traditional do-it-yourself (DIY) networks and focus on core business objectives.
Furthermore, NaaS enhances user experience through high network performance levels. It ensures network performance meets the required standards, providing end-users with a seamless and efficient connectivity experience.
Nile’s NaaS security has security by design, and the Zero Trust approach simplifies and enhances protection against advanced threats.
My take on Nile’s NaaS offering
Ok, I like Nile’s solution. Their technology is amazing, their operational model is impressive, and how they run the network in a zero-trust manner is highly secure.
There is a market where customers would like to outsource everything in networking to a 3rd party.
However, is it for everyone?
Let’s discuss if we can answer this!
The NaaS concept is inspired by the cloud concept. What cloud has done to computing, NaaS has the potential to do to networking.
Today any cloud offers a variation of offerings from IaaS to PaaS and SaaS
As you may understand, the success of the cloud is in offering a “variation” rather than a one-size-fits-all offering.
A company deciding between IaaS, PaaS, and SaaS depends on multiple factors; for example, the skillsets it has to run and maintain an application.
In other words:
A company that does not want to run /maintain a virtual machine or an application will go with a SaaS offering; such a company does not care about the underlying infrastructure ( whether it runs on AWS or Azure etc.)
On the other hand, a company that has the resources and skillset to deploy and maintain the infrastructure may go for an IaaS or PaaS.
There is no absolute wrong or right way here. This will depend on the company’s skillsets and resources to decide which way to go.
To me, the Nile offering combines good products and good service. However, the only way to get the product is as a “total solution,” i.e., a Network-as-a-Service offering.
I am trying to draw NaaS in parallel with cloud offerings such as IaaS, PaaS, and SaaS.
Comparing it with the cloud, NaaS is more like a SaaS offering. That means the underlying infrastructure is like a black box for a customer.
To many customers, it is fine to consider the network as a black box. To a lot of others, it may not be what they want. For example, medium to large enterprises already have the skillset and resources to run their networks. Thus they would like full or partial control over the network instead of a NaaS offering.
Effectively Nile may miss a broader market that needs part of their solution.
For example:
A customer may be interested in the technology but not the operation. They have the engineers to run the network efficiently.
Some customers have the network already but would like to explore the security offering from Nile.
Some would like to integrate their networks with the Nile portal for management to take advantage of Nile’s management and analytics.
Nile uses sensors to collect data from the network, which helps them in smart troubleshooting. There would be an interest in the market to buy that part only.
Can Nile address the different segments as above?
No!
OK, strictly, the above market is not a NaaS, which is not what Nile stands for!
However, Nile can package a variation of their NaaS offering; they can address a boarder market. Nile can keep the As-a-service part but more modular and part offering rather than a total NaaS.
This would require Nile to think out of the box and package its NaaS offering for different market segments in different ways.
This would also eliminate the fear of vendor lock-in. Buying end-to-end NaaS solution from one vendor.
In short, Cloud was successful not because of a one-size-fits-all offering but multiple options, whether SaaS, PaaS, or IaaS, that suit different customer segments with different skill sets. There is room to bring that variation in NaaS as well.
Disclaimer: This is my own analysis. The vendor did not ask for nor were they promised any consideration in the writing of this post. This is not a sponsored post. My conclusions here represent my own thoughts and opinions.
I had a chance to watch an interesting presentation by Kentik in NFD-31 online back in April 2023. I have been attending NFD events organized by Techfieldday events for quite some time as a delegate. However, this time around, I just watched the presentation as a viewer ( not a delegate)
Kentik has been coming regularly to these events. The last one where I was a delegate was the NFD service provider two in August 2022, and Kentik presented there too.
For a while, I wanted to write about Kentik considering their unique approach to “network observability.” The presentation in NFD-31 by Phil Gervasi about “Data-driven network visibility with Kentik” clearly and well-explained how Kentik solves the network observability challenge. Thus it became a good inspiration for me to write a piece about Kentik.
For those who do not know about Kentik:
Kentik solves network operations problems using a data-driven approach called network observability.
Kentik Network Observability Cloud platform is their flagship SaaS solution. Which analyzes a large amount of data from various sources and uses machine learning techniques to classify, cluster, scale, and normalize the data. The goal is to provide useful and insightful information to help network engineers resolve issues quickly.
This approach differs from point solutions in the market, providing a narrow network view. More on it in a moment.
Before we discuss further, we should understand the industry drivers:
Do you know what is the biggest challenge of enterprises today?
No matter how important the network is today, they do not own it.
Really?
Yes, with the cloudification of applications, the services now reside in private or public clouds. Many applications have moved to SaaS.
Network owners have become dependent on networks beyond their premises and their control.
As enterprises march towards digital transformation to provide new and innovative services and strive to improve the customer experience, they must ensure that the network ( even if they do not own it ) and applications perform well.
This was quite different in the old days.
Earlier, finding the issues was easy; the network was within an enterprise’s control. Now the network has many pieces, the internal network, the public network, SaaS, security devices, load balancers, DNS, IPAM, containers, etc.
Getting visibility across all of them is indispensable. At the same troubleshooting has become complex.
The traditional monitoring and testing tools do not suffice here. The monitoring tools give a partial view of the network. This would require multiple tools to get end to end view of the network.
There is a lot of data. However, correlating them to find the root cause is a challenge. Numerous disparate databases and point tools further slow the network management and troubleshooting process.
This requires “observability,” as Kentik calls it. Observability runs at the heart of Kentik tools.
Network observability provides a greater level of visibility and context compared to traditional visibility. It involves using tools and techniques, such as data science and database architecture decisions, to understand why something is happening in the network.
Observability does not involve making changes to the system but observing and analyzing production and user traffic. Additionally, active monitoring tests the network without relying on production traffic or user complaints.
Fig: Traditionally, visibility vs. Network observability – ( Ref. Kentik)
How Kentik is different?
The following diagram nicely summarizes what Kentik does.
The strength of the Kentik is first in collecting the data from many diverse sources, which would otherwise require multiple tools. This includes the NETWORK DATA sources such as routers and switches and APP DATA like CDNs, containers, hypervisors, and servers. And if that is not enough, they can take BUSINESS DATA also like CRM, OSS, BSS, etc.
The data goes to a big repository where data is ingested and fused. In other words, data received from customers is appended with additional data such as GeoIP and BGP, which gives additional context data.
Fig: Kentik’s approach to observability (Ref. Kentik)
However, Kentik goes one step further; it analyzes and correlates multiple data sources to provide meaningful insights into what is going on in the network.
It presents charts and statistics about anything related to the network and application it collected, from which pinpointing the exact problem becomes easier.
It can help answer questions like
Why application is slow?
Which ASN is adding to the latency?
Why a page is taking longer to load?
Is application degradation came because of Jitter?
Last but not the least? Is it a network issue or an application issue?
The synthetic test is one of the nice features of the tool, through which traffic can be inserted into the network to see the performance of test traffic, thus simulating user experience.
Phil explained the functionalities of the Test Control Center, highlighting the ability to identify and troubleshoot issues before they impact end users. He gave an example of using a Page Load Test to identify high latency and traced it back to a specific file causing loading delays. The historical data can be shown in time series, as in the following picture,, and periods when the user experience went bad with the page load.
Phil showed how easy it is to drill down using the waterfall method for a particular .js file showing issues. It was found that the file seemingly had trouble resolving DNS ( DNS load time of around 6 seconds), resulting in an overall slowdown of the application. Phil emphasized that such issues could be fixed before users experience them, ensuring a smooth running system.
Fig: The root cause of slow page load – Ref- Kentik
What are the advantages for enterprises and Service Providers?
As you may have seen, Kentik removes the guesswork in network troubleshooting.
You do not need tens of engineers and tens of tools to find out issues in the network.
Even if you have tens of tools, arriving at the root cause of the solution is lengthy., time-consuming, and manual.
Digital transformation is all about services, agility, and customer experience. When customer experience matters, time is of the essence.
Kentik fills visibility gaps and prevents disruptions in internal networks and cloud environments. It helps network owners stay agile and proactive. Kentik Network Observability Cloud empowers professionals to plan, run, and fix any network. With this platform, they can easily identify where the traffic was impacted and take corrective actions. Reliable networks and a great digital experience are ensured.
Disclaimer: This is my own analysis. The vendor did not ask for nor were they promised any consideration in the writing of this post. This is not a sponsored post. My conclusions here represent my own thoughts and opinions.
“Disaggregated Networking” is good, but “Network Cloud” is better.
This is my analysis for one vendor’s approach to routing, after attending the NFD event on service provider that was held December 8-9 and that was Drivenets presentations on network cloud. I was one of the analysts that attended the event.
Disclaimer: This is my own analysis. The vendor did not ask for nor were they promised any kind of consideration in the writing of this post. This is not a sponsored post. My conclusions here represent my own thoughts and opinions.
Coming back to the routing in general, for a while I wanted to write that routing needs a fresh approach.
Routing Needs New Approach
During the last twenty to thirty years, we have not seen a lot of innovation in routing. For example, we still see the router as a monolithic piece of software tied to the hardware. The software is “made for the hardware”. New features in software highly depend on the hardware itself. In many cases, it means an upgrade of the hardware to support new features.
This tight integration means that routing is mainly dominated by a few big vendors. This is a good business after all to be in both hardware and software as a business in one can translate into business in another area.
This is not good for the end-user, though. For the end-user, it means sometimes waiting for months or years to get new features as the vendor needs to integrate and test any new feature with its own hardware or develop that “new card” as the current one is not compatible. Needless to say, it’s a barrier to the time to market to and service innovation.
Is basic Disaggregated Networking an answer?
Yes, but not enough!
Arguably, this should be the first approach to the routing.
“Disaggregated networking means software is abstracted from the hardware. We can buy a white box from one vendor that runs a commodity ASIC such as Broadcom and software, from another vendor”
Fig: Disaggregated Networking
On the other hand, this approach is still not enough.
The software could still be monolithic, albeit now we can buy the software from one vendor and the hardware from another, so we have more choices.
From “Disaggregated Networking” to “Network Cloud”
While Disaggregation is the first step. Taking it to the next step of “Network Cloud” opens the door of innovation.
They have approached networking with the cloud perspective, thus taking it from disaggregated networking to “network cloud”
Come to think of it! We have seen server innovation through virtualization and cloudification over the years. We see obvious benefits of the cloud when it comes to compute such as scale in and scale out of resources. Above all, cloud-native applications take it one step further. With cloud-native applications, it is easy to innovate, add new features, and scale the apps in a breeze.
This is also disaggregation but at a different level.
Drivenets has translated that into networking and what they call it “Network Cloud”
Consider the diagram below:
Fig: Network Cloud
With “Network Cloud”, the white box layer becomes a resource pool. It is treated as a cluster. A network hypervisor is introduced that abstracts the networking resources.
This is augmented with a cloud-native apps layer. With this, the cluster turns agile. It can work as an edge, aggregation, or core routing layer. security apps like Firewall and DPI can also be introduced. Or different apps can be run in mix environment.
Now think of it. The hardware layer is not changing at all. While all the innovation is at the software layer. The hardware is just a resource that can be scaled separately.
All this with the benefit of the cloud. Start with a single white box, then start adding more white boxes in the cluster as the expansion needs arise. This provides a “scale-out” approach as promised by the cloud. Not to forget that the approach is more resilient as it does not tie the software to any particular box, so a failure of a box does not affect any application.
I believe it is time the approach to routing is changed. The Mobile core has seen it with the cloud-native 5G core, the transport needs this now. The legacy and monolithic way of building the routers impede service innovation. The cloud-native way of routing cuts down on applications development time and provides a more agile way of launching new services. The service providers must look for out of box approach if they want to evolve their routing layer quickly and enable it for service innovation rather than treat it as a cost center.
I was invited as an analyst to attend the NFD-26 event organized by Tech field day featuring networking vendors.
As I listened to different vendors’ sessions, one of the presentations that caught my interest was from Juniper regarding their vision on AIOps in networking.
As I work for a service provider, I deliberated on what issues this can solve for a managed service provider (MSP)
First, the services have moved out of data centers/enterprise premises to the cloud; this decentralization of services means multiple diverse networks domains exist between the user and services i.e. wireless, wired, and WAN.
More networks mean, more complex management and difficulty in troubleshooting, should anything go wrong.
For example, a zoom call starts from Wi-Fi, but crosses the wired network and then WAN to reach Zoom servers. Therefore, any issue in any domain can result in poor call quality. This demands a consistent and comprehensive way to troubleshoot, service from its origin to its destination.
Second, the users/enterprises have become more experience-centric which means they are more conscious and demanding from their service providers on the consistent quality of experience. Therefore, MSPs must find ways to find and present the actual user experience to their customers.
Need for AIOps rather than conventional Ops
Let’s face it. the networks are far more complex and dispersed today.
The diverse networks, today, generate a lot of data that the conventional Ops tools are not able to handle. An issue that affects a user experience could be a wireless issue. But it might very well be a switch port issue or it may be an issue in the WAN.
Having 360-degree end-to-end network visibility and effective troubleshooting across wireless, wired, and WAN require an altogether different approach during Day 2 operation.
Automating network operations is the key to expedite fault detection and resolution.
Could AIOps come to our rescue?
Perhaps Yes,
But then how many tools in the market offer these capabilities?
Juniper Mist AI fills the gap of the tools needed for end to end network AIOps
As I attended the session on Mist AI from Juniper, I realized that this could be a solution to the problem MSPs face today.
As you can see in the figure, their target is to have end to end AIoPS which is not limited to Just wifi networks but also wired and WAN networks
But that is not all.
Mist AI focuses on the “User Experience”
AI Mist focuses on providing visibility into the real” user experience”. So rather than a typical dashboard that provides network KPIs and faults overview, they go few steps beyond:
Mist AI can collect data from thousands of wireless APs, wired switches, and WAN networks ( Juniper has integrated their SD-WAN routers with Mist AI) and drill down any issues that could affect a single client.
What impressed me was the conversational interface which makes the tool very user-friendly. A user can ask questions like ” What is the issue with Bob’s connection today?” and it can return answers like ” Bob has a wireless issue and here are the steps on how to solve them” The user can ask further questions to drill down and the AI assistant can provider further actionable result and feedback.
Combined with the Mist AI WAN assurance app, the AI tool can provide insights into WAN. This is quite powerful to bring data from multiple different domains including wireless, wired, and WAN under one dashboard. Which in turn enables proactive actions, automated workflows, and actionable insights into network issues from users to the cloud.
Conclusion
In short, Mist AI seems very powerful and brings a different approach to the way Ops can approach the network. Juniper has brought a fresh approach to the way Ops should work today. For the MSP, it can cut on the time of troubleshooting and enable an experience-centric service to its customers rather than a network-centric experience.
Disclaimer:
Juniper was a presenter during Networking Field Day 26, a virtual event organized by Tech Field Day ( Sept14-16). I was invited as an analyst to listen to different vendors’ presentations. Vendors did not ask for nor were they promised any kind of consideration in the writing of this post. This is not a sponsored post. My conclusions here represent my own thoughts and opinions.
So it is not easy to have a list of OpenStack components/services all in one place. They are everywhere on the web and one can easily get lost. and when you do see a list of services, they are not described in a simple way.
So here is an easy-to-understand list describing the OpenStack components for a quick reference. Whether public clouds or private clouds, OpenStack is everywhere, so you better understand these terms.
What you see in this diagram is the list of core services already described in my blog here which but there are many others, it is an attempt to cover as many OpenStack components and services as possible in this list.
Nova manages pools of compute resources including VMs, bare metal (through use of ironic), and containers (limited support). The good thing about Nova is that it is hypervisor agnostic so it can use KVM, VMware, LXC, XenServer, etc.
As more and more apps are being containerized, cloud users need a more direct way to manage containers. Zun provides a simple way to manage containers through APIs from within OpenStack without having the users understand the complexities of containers. While Nova needs a Nova docker driver for containers management, Zun can do it without depending on the Nova API.
It is a bare metal provisioning service. Although OpenStack generally provisioning Virtual Machines (VMs) but sometimes it is required to provision bare metal. This service allows this through the use of APIs from within OpenStack. Also, it integrates with Nova, enabling Nova to use its services for provisioning bare metals.
The need for Telco workloads requires the use of accelerators ( GPU, FPGA, ASIC, NP, SoCs, NVMe/NOF SSDs, ODP, DPDK/SPDK). Cyborg creates and manages accelerators with either tools or the API directly.
Cinder is an OpenStack Block storage service that allows you to add persistent storage to your virtual machines. It interacts with OpenStack Compute to provide volumes for instances and enables management of volume snapshots and volume types.
Have you heard of S3 in AWS which is a cheap way of storing a huge amount of data? Swift provides the same service. It is also called “Object Storage”. It is highly distributed and provides a cost-effective scale-out object storage
Magnum takes a step beyond Zun. While Zun launches and manages containers, Magnum makes the more complex orchestration engines like Kubernetes or Docker Swarm, available as a resource in OpenStack. So if the intent is to manage the containers’ engines ( rather than containers), then Magnum is the option.
With Neutron, users can create networks and connect devices/servers and services. They can manage IP addresses, create subnets, VLANs, private network, etc. to manage such communication. It provides networking as a service.
Octavia provides load balancing services to virtual machines, containers, or bare metal servers. This is done on demand. This on-demand horizontal scaling is a differentiating feature for Octavia compared to other load balancing solutions
Keystone enables authorization and authentication and thus is an OpenStack identity service. It is used to query which users are authorized to use a cloud service. User name and password credentials are some of the methods that the Keystone unit supports. It is possible to integrate it with systems like LDAP.
Users can discover, register and retrieve virtual machine images using Glance. It offers a REST API that enables you to query virtual machine image metadata and retrieve an actual image. You can keep virtual machine images in a variety of locations, from simple file systems to object-storage systems.
In order to help other services effectively manage and allocate their resources, Placement is an OpenStack service that provides an HTTPAPI for tracking cloud resource inventory and usages.
A new project called Senlin provides a generic clustering service for OpenStack clouds. While Heat can provide such services, but it was decided to offload such function from Heat and provide a dedicated clustering service that can help in auto-scaling.
The aim is to bring big data and OpenStack together. Sahara enables creation and management of Hadoop cluster ( or Spark), all from within OpenStack without the need for dealing with another cluster management app
Trove is a database as a service that runs on OpenStack. It‘s designed to allow users to quickly and easily use the features of a relational database without having to deal with complex administrative tasks. Users and database administrators can provision multiple databases as needed.
Introducing an application catalog to Openstack, the Murano Project enables application developers and cloud administrators, to publish various cloud-ready applications in a browsable categorized catalog. The actual deployment is done by the orchestration tool such as Heat.
A freezer is a distributed backup and restores as a service platform that enables Disaster as a service in OpenStack. It is designed to be multi-OS ( Linux, windows, etc).
OpenStack Dashboard (horizon) provides administrators and end-users with a graphical interface for accessing, provisioning, and automating the deployment of cloud-based services.
How you should plan? where to place the controller? and where to place the rest of the functions?
These questions are the first ones that pop up when a Telco wants to design its edge network
After all, the reason a telco decides for edge computing in addition to cloud computing is to run real-time functions closer to the user, so it is important to focus on the location of functions placement in the edge cloud architecture.
It boils down to where the control functions are placed in edge computing architectures.
In addition, the edge architecture should be flexible enough.
In other words, the architecture should not tie to only ONE specific environment such as containers or VMs or even bare-metal.
Rather, it should be flexible enough to run any environment. Also, the architecture should equally apply if an operator wants to use a public cloud or a private cloud.
With these factors in mind, there are two popular architecture models which I will explain, but to understand them, it is important to understand the difference between the central data centers and the edge data centers, first.
Central Data Centers vs Edge Data Centers
Central data centers ( central cloud) are the usual ones we know of, that are more centralized and far from users, while the Edge Data Centers are located closer to the edge as shown in the following diagram from Openstack
Centralized cloud runs control plane or signaling plane-related functions. Examples from core networks include the signaling plane of IMS or the control plane of EPC.
While the distributed Edge Data Centers run mainly “user” functions. These are the functions that are throughput intensive and are latency-sensitive ( also called real-time) so they should be run as closer to the user as possible. One example of a user function is “UPF” in 5G. In terms of the actual applications that are suitable to be run on edge network are video surveillance, CDNs, AR/VR, etc.
Also in the C-RAN approach, CUs and DUs can be run closer to the tower at the edge data centers.
So in summary, the general principle of deciding what to run in the central layer vs edge layer is following
If the workload is control plane/signaling intensive, it is run at the central location
If the workload is throughput intensive and/or latency-sensitive, it is run from the edge location.
Central Data centers vs Edge Data Centers Ref: Openstack
Two Models of Edge cloud architecture
While there may be a lot of edge computing architecture models, the following two are the popular ones.
Centralized Control Plane
For the centralized control plane model, the controller is placed in the central data center while the edge data centers carry the compute nodes only, sometimes also called the edge server.
The controller includes the following functions
orchestration
authentication
storage management
image management
Centralized Control Plan-Edge Cloud Architecture
The management and orchestration are done centrally, so the advantage is the “ease of control” from a central location. On the downside though, the edge data center can get isolated if there is a loss of communication between the two data centers resulting in the edge data center running without any control.
Secondly for such a model to work effectively, there needs to be a good connectivity layer between the edge data centers and centralized data centers as there is a lot of dependency on the effectiveness of the connectivity
Therefore this takes us to another design of edge cloud architecture
Distributed Control Plane
In distributed control plane architecture, the controller instance runs on every edge data center thus making it autonomous. In other words, the edge server runs both the control functions as well as the compute functions.
There are a couple of ways to operate such a model. One way is to create a federation of edge data centers and connect their databases to operate infrastructure end to end as a whole, another way is to synchronize the databases across sites to have consistent configuration across databases in the edge data centers
This model provides more resilience as having a loss of communication between the central data center and edge data center would have less impact. As the configurations required to run the edge data center are managed locally.
Distributed Control Plane-Edge Cloud Architecture
Note that these architectures would apply also to the public clouds. For example, a telco may decide to use a public cloud at its edge node instead of its own cloud. In that case, both the above design options are also applicable.
In summary, designing a network edge is not a random precise but deliberate attention should be paid to the choice of architectures available.
Some choices that can affect your decision are following
1. The choice of deployment depends upon your requirements for scalability, resiliency, cost-effectiveness etc.
2. Do you have sufficient space in edge data centers to accommodate control functions placement?
3. What level of redundancy do you need for network edge?
4. What are the geographical distances? if the distances are small, the centralized model is easy to run and manage.
and “what is the latency threshold for each edge data center” is the second that comes to mind
Therefore, this blog answers these questions in easy to understand manner.
I do see a lot of confusion with the naming and the latency when it comes to edge computing
For example, there are as many names for edge data centers as there are types. We haven’t heard such names for the traditional data center:
Near-edge, far-edge, deep-edge…..etc.. you name it and you will find the type
Don’t get confused by them. A lot of them are just the marketing terms coming from the edge computing vendors. So really does not matter.
In fact, Edge can be anywhere ( between centralized data centers and users), it can be at the customer site, at a base station site, one hop away, two hops away, etc
However, I do think that everyone needs some reference and benchmark.
So I will use two sources: OPNFV and OpenStack foundation and show latencies and locations from these sources
Just one point, the E2E latency is one way and not two way and I have explained it well in this article on how many people get confused with one way and two-way latency.
The official documentation for OPNFV defines E2E delay as following
“The time of the transmission process between the user equipment and the edge cloud site. It contains four parts: time of radio transmission, time of optical fiber transmission, time of GW forwarding, and time of VM forwarding”
Small edge data center is closest to the base station.
Latency from UE to the site is around 2 ms
Distance from cell site to edge site is around 10 KM
The maximum bandwidth it can provide is 50Gb/s
Services that can be deployed here have very low latency requirements.
It is possible to use this site without virtualization using just bare-metal
Medium Edge / Aggregation Edge
This is larger than access edge and can also serve as an aggregation site of multiple access sites.
Latency from UE to the site is around 2.5 ms
Distance from cell site to edge site is around 50 KM
The maximum bandwidth it can provide is 100Gb/s
It is possible to use bare-metal as well as VMs/containers
Large Edge/Aggregation Edge
This has a larger capacity than the medium edge and further way at a distance of 80 to 300 KM, this can also serve as an aggregation layer.
Latency from UE to the site is around 4 ms
Distance from cell site to edge site is around 80-300 KM
The maximum bandwidth it can provide is 200Gb/s
It is always virtualized using VMs or containers or both
What are the other locations of edge data centers?
Edge data centers can also be located right at the base station, called tower edge, or on the customer premises, the on-premises data center are owned by the enterprises themselves.
What is the best location for an edge data center?
It depends on the latency requirements and the need for reducing the backhaul bandwidth. There is no one-size-fits-all location. It can be within the premises of the customer, at the tower, one hop away or two hops away depending on the application latency requirements.
For telco edge cases, some guidance can be obtained from OPNFV here
CDN – 10 ms
enterprise vCPE -50 ms
5G-UPF 10 ms
MEC uRLLC < 3 ms
eMBB < 10 ms
CRAN-CU 3ms
Note that the CDN’s ( content delivery network) latency is not that strict and the motivation of its deployment may be mainly reducing the backhaul transport rather than reduction of latency.
Edge computing is not only about having technology but also a smart business model.
This statement is particularly true for telcos.
And the telcos know about it. They know that with the new 5G core, they have an opportunity to deploy a more intelligent and service-oriented network edge which can open new business avenues like ultra-low latency use cases (AR/VR, connected cars, IIOT, low latency gaming, etc.)
However, they also know that they need have a proper go-to-market strategy (GTM)
Without a smart edge business model, they may leave the “greenfields” to more active players like the hyperscalers/public cloud players (AWS, Azure, Google Cloud)
Yes, the cloud players are quite active and aggressive in the edge area. Amazon’s “AWS Wavelength“, “Azure Edge Zones” and Google’s “Anthos for Telecom” are just a few examples that show the ambitions of these players.
On the other hand, the telcos/CSPs are not completely disadvantaged either:
They own the network-which is their biggest asset. Therefore, they can use their base stations, pops, and central offices to bring computing resources as closer to the users as they like.
Why this is important?
The edge applications can be “edge-aware” or “edge unaware”. However “edge-aware” applications are becoming more important and popular and they need to communicate with the network and understand network information like latency and location. In short, application owners need to work closely with the telcos in this area.
No doubt, hyperscalers have a stronghold when it comes to the edge applications marketplace.
But in the edge business, they will still need to collaborate with telcos closely.
But how they can work together? Why not a telco introduce a marketplace of its own?
Unfortunately, there is no ONE answer to these questions?
So with this piece of article, I want to write about the different business models telcos may have in relation to hyperscalers in the area of edge business. The different ways in which they can compete as well as cooperate.
Nevertheless, when it comes to the field of Multi-Access Edge computing (MEC), telcos need to play smart with cloud edge to synergize with the hyperscalers.
Let’s explore the different Edge Cloud business and operational models for telcos. This is purely a comparison of operators and hyperscalers for the telco edge. There may be different business models for telcos other than that.
However, I would like to pause the discussion for a moment, and have a quick refresher on the high-level view of the edge stack, as this is important to understand before understanding the business models.
Fair?
Let’s dive in.
What is the Edge stack? Understanding it before understanding edge computing business models for telcos.
The following figure is adapted from LFedge that shows multiple layers.
Edge Stack for Telco Edge
Operator’s Infrastructure
This is the “access network” used for Edge services. This can be a fixed network like FTTx or it can be a mobile network used for Telco edge. The edge location can be anywhere like a base station site, one hop away, two hops away or further depending on the latency requirements. Anything related to the operator’s infrastructure is included here like 5G core but in addition, may also include virtual resources ( VMs) so that edge applications can be onboarded.
Edge Enabler
The Edge Enabler software layer sits on top of the operator’s infrastructure. It is the layer closer to the network and as such has the network information like radio quality, location awareness (important for “edge-aware” applications). In the ETSI MEC architecture which I explained here , this is the “MEC platform” that acts as an Edge enabler. Normally such a platform is provided by the telco as Edge enabler interacts with the 5G core network using 3GPP APIs. But a Telco may work on a model to have a third party bring its MEC platform and use only its infrastructure.
Application Enabler
Bringing Edge Enabler is not enough, there needs to be an application enabler too. Application Enabler is like an abstraction layer on top of Edge Enabler.
For example, do you think the application developers for edge apps want to bother about 3GPP interfaces and how they work?
They do not care about it. Application Enablers sits in between Edge Enabler and edge applications Its mission is to provide edge app developers friendly APIs, allowing them to consume and manage specific telco network capabilities without having to know about the underlying telco network!
Edge End-user App
This is self-explanatory ( i.e. end-user edge application)
Edge computing Business Models for telcos
Let’s take into account the different operational models for telcos in the area of edge infrastructure.
Where there may be multiple models for telcos to work and there may other edge service providers besides the hyperscalers, but I am focusing on hyperscalers here to make the demarcation between the two clearer.
This can be called the “Commodity Telco” because it is present at the bottom of the business value chain.
The Telco does not have any of its edge services, neither own its MEC platform. It just provides its infrastructure for example 5G core. The hyperscaler brings its Edge Enabler layer (MEC platform) and integrates it directly with the infrastructure. The rest of the stack all belongs to the hyperscaler including the edge marketplace portal.
As the marketplace is facilitated and provided by the hyperscaler through its portal. The app developers have a direct interface with the hyperscalers with no visibility for the Telco.
Most likely the hyperscaler is paying for using the operator’s infrastructure on some basic usage model such as a “pipe model”
We can say that the Telco here is completely “commoditized” and as such not a recommended approach for telcos.
The Edge enabler Telco-Option B
In Option B, the telco introduces its MEC platform as Edge Enabler. While the hyperscaler brings its application enabler platform along with its marketplace.
This can create a win-win situation for the telco. The telco has risen in the value chain above. The ROI for the Telco is higher compared to the commodity model in Option A and can charge a premium service for providing its MEC platform.
In this model, the marketplace is still owned and controlled by the hyperscaler. However, the telco may work out a revenue-sharing model with the hyperscaler instead of offering flat fees or a basic consumption model.
The Full Edge Telco-Option C
This is the easiest to implement option for the Telco
In this Option, the Telco owns and operates the complete stack.
The telco can provide a vertically integrated service to its customers. An example service is a private edge service for its enterprise customers such as industrial IoT. A variation would be offering services up to the application enabler layer and let the enterprise customers use their applications.
The telco is free to price and charge its customers, the way it likes.
The Dream Telco-Option D
This is the Dream Telco because every telco would like to be in this place although most complex to implement.
As compared to Option C, the telco now offers its marketplace so that developers can develop the applications and sell them on the marketplace. The telco has full control of the marketplace and can charge a premium for its services. The revenue potentials of the Telco are much higher as the Telco is working now in the “applications” space and at the top of the value chain.
Recommended Approach for Telcos.
Unfortunately, there is no one best option considering the current state of telcos.
Option A is NOT recommended
The Commodity Telco such as Option A is NOT a recommended approach. A passive telco with no edge strategy risks being approached by a hyperscaler offering a commodity model. The telco has no visibility to applications, not getting any premium revenue here.
Options B, C, D are recommended approaches
In my view, the recommended approach for telco is a mix of Option B, C, and D. ( that is either choosing one of them or combining them)
If a telco would choose, of course, it would like to be the Dream Telco.
However, the Dream Telco needs to run a full stack including its marketplace. To be successful, it needs to have DevOps of its own and an ecosystem of developers, and a very aggressive go-to-market GTM strategy. As it stands, the majority of Telcos are at the early stages in all these three.
However, there are still a handful of Telcos that can fit in this category but not the majority.
Conclusion: Starting with Option B and C together is the best
For the considerable majority, starting with options B or C in parallel would be practical and easier to implement. They need to join hands with the hyperscalers and in parallel have a full edge stack for its private customers. This would give them quick access to the market places which the hyperscalers already own.
Come to think of it, collaborating with hyperscalers gives easy to access considering the maturity of the ecosystem they provide. Any developer today would like to integrate with Azure, Amazon, or google cloud. To reach that level will take many years for a lot of telcos.
To get the benefit of the edge today, that seems to be the practical option to start collaborating with the hyperscalers.
However, in the long run, the telcos can aim for Option D to be completely independent and run its edge network on their terms. For this to happen, the telco needs to have a strategic plan of digital transformation at all layers: technology, organization as well as business layers.
So do you agree with me or not?
So what is your tack on telcos vs hyperscalers in the telco edge area?
Can telcos compete with the AWS wavelength and Azure edge zone, if they remain in the status quo?
Is Edge computing the only hope for service providers?
With billions spent on cellular networks, everyone asks about the killer use cases for 5G.
The new 5G Core is capable to provide ultra-low latency services opening the network to potential new user cases, but it can not solve on issue-The distance and latency.
However, this is not the issue of the 5G core either. This relates to physics where latency increases with distance.
That is where edge computing can come to the rescue.
Instead of bringing the traffic to the centralized cloud (in central data centers), which are far, bring the cloud resources to the user itself.
Edge computing brings intelligence to the network edge ( where edge infrastructure is deployed) and thus has shown the potential to open the network for the killer use cases for 5G like autonomous vehicles, connected cars, AR/VR, remote surgery, etc. Also, the increased use of edge devices because of IoT has resulted in a need to take real-time action on the data closer to the network edge.
There are a lot of articles on the internet on “what is Edge computing” but this article is different. As always the case I clarify the concepts by comparison. First, this is your one-stop-shop for understanding edge computing. Secondly, I will clarify the concepts by comparing them with cloud computing, further, also clarify by comparing “edge computing” with Multi-Access Edge Computing (MEC).
So stay tuned till the end.
Just one point though, I want to hold the discussion of MEC ( and its difference with edge computing), until I reach the point where I discuss the location of edge cloud as it will clarify the concept to you at that point. However, in one sentence, MEC is a subset of edge computing so the concepts apply equally.
Fair enough? Let’s proceed.
What is Edge Computing and MEC ? (Definition)
We are living in a connected world, With the enormous growth of connected devices, and the demand for ultra-low latency services, the proliferation of mobile devices, edge computing has picked up considerable momentum.
One of the signs of the growing interest is to see how many organizations are involved in the standards. There are many of them in edge computing: ETSI, 3GPP, Linux Foundation, GSMA to name a few, etc.
While each industry body has defined the word “Edge computing” I personally like the one form LF Edge body under Linux Foundation, which is short and comprehensive
“Edge computing represents a new paradigm in which compute and storage are located at the edge of the network, as close as both necessary and feasible to the location where data is generated and consumed, and where actions are taken in the physical world”
For context purposes, I will let me all also add the definition of MEC by ETSI as they own the standardization of MEC, but I will delay the discussion of the “Edge computing vs MEC” to the end as it is important to grasp some initial concepts first.
“Multi-access Edge Computing (MEC) offers application developers and content providers cloud-computing capabilities and an IT service environment at the edge of the network. This environment is characterized by ultra-low latency and high bandwidth as well as real-time access to radio network information that can be leveraged by applications”
So what is Edge computing in in simple terms ?
In simple words, edge computing refers to running “cloud” closer to network Edge. And “close to the Edge” means closer to the user.
But why would one run the cloud closer to the user in a decentralized way?
What are benefits of edge computing?
Reduced latency is the first and foremost benefit. Applications need to run closer to the user to improve the user experience.
But that is not the only benefit.
Reducing backhaul bandwidth is the second most important benefit.
With so many applications running at the edge, many of them, for example, video, sending the data all the way to the centralized cloud is bandwidth-intensive and hence CAPEX intensive too.
Why not just process them at the edge in a decentralized way, thus saving on that bandwidth ? and hence costs.
Hold on! Did someone say “decentralized”?
Isn’t it counter-intuitive to the concept of “cloud”?
Because when we think about the cloud we think of pooling the resources at a central place so everyone can use them.
On the other hand, Edge computing means resources are de-centralized.
So it is the right time to understand how Edge computing is different than cloud computing?
How is “Edge Computing” different than the “Cloud Computing” ?
To understand the difference consider the diagram at the left. All the processing and decisions on the data are done at the central place in the cloud computing case while in the edge computing case they are done closer to the user in the Edge cloud ( many clouds distributed).
You must note, HOWEVER, that edge computing is not a replacement for cloud computing. Not all applications run from the edge. Where latency requirements are not high and bandwidth requirements are low, it is always good to run these applications from a centralized cloud to take the advantage of resource pooling.
Also, many applications will need both cloud computing and edge computing. With user-intensive traffic processed locally at the edge and control traffic sent to the centralized cloud.
Therefore although edge cloud would not replace cloud computing, it will certainly reduce the need for it for some applications so it can be scaled down.
Centralized cloud vs Edge cloud
Let’s try to compare them on various characteristics.
Feature
Edge Computing ( distributed )
Cloud Computing ( Centralized)
Compute
Medium to Low computational capacity
High capacity compute resources placed at a central location
Latency
Reduced latency as the application is processed in edge
High latency. It depends on how far is the centralized cloud
Real-Time data processing
Best for real-time data processing
Does not provide as good results as Edge cloud
Backhaul Requirements
Very less as data is served from the edge so no need to see all data to centralized cloud
High backhaul requirements
Security Requirements
Higher security requirements as cloud surface is more
High-security requirements
Best for what kind of applications
Latency sensitive Or/and Bandwidth intensive
Latency agnostic and the bandwidth needs are low to moderate
Edge Computing vs Cloud Computing
After knowing the difference between Edge computing and cloud computing, it is time to move to understand the characteristics of Edge computing Characteristics/ Attributes of Edge Computing
So it should have the basic attributes of cloud computing, but it should offer more because of the distributed nature of the edge computing
The following are the important key attributes of Edge Computing according to ITU-T Y.3500 and this research paper. I am summarizing them for your easy reference
Now did you notice something?
While the latency and bandwidth are clear benefits that edge computing provides and discussed earlier but additionally edge computing offers network information and location awareness because of the close proximity to the network. This is important as the edge site is near to the access network, an application can take real-time action based on the radio network information ( in the case of the radio access network thus helping real-time applications)
Traditional Cloud Computing Attributes
Additional attributes because of the Edge nature
Broad network access like 5G/LTE, FTTH, WiFi
Multi-tenancySelf-service- Ability to provision without the involvement of cloud service provider On-demand
Rapid elastic and Scalable
Resource Pooling
Measurable Usage
Low latency
High Bandwidth- Provides high Bandwidth as processing happens locally
Real-time insight into Network information ( radio network for example) which enables taking action based on radio statistics ( applies to wireless access networks)
Location awareness- Can be aware of users locations by analyzing the information received from user devices
Cloud computing vs Edge computing
Where is Edge ? ( Location of Edge Cloud)
There are two distinct categories of Edge Data Center, the User Edge, and the Service Provider Edge
OK don’t get confused with the names, there are many types of names, there is deep Edge, Far Edge, Regional Edge, enterprise edge, etc..They may all be the marketing terms of vendors
However, as a starting point, I will say focus on two areas here. The service provider Edge and the User Edge. This is taken from LFEdge ( Linux foundation)
Service Provider Edge is where the network of the service providers is present while the User Edge is at the premises of the user. Now it does not mean that the service provider cannot manage the user Edge. In most cases, this will be the case, if a service provider establishes a DC within the premises of the customer OR puts a CPE inside the data center of the customer ( on-prem. edge computing solutions).
But broadly speaking, Service Provider Edge is the edge of the service provider network. A service provider may need to establish a small pop for the edge data center providing similar components as a traditional data center.
On each of these edge locations, you can put edge servers ( edge servers provide compute power)
Did you notice something?
If you look at the Service Provider Edge. The green color is extended inside the User Edge, which means that the boundaries of the service provider edge are not strictly defined. and it can go inside the customer premises also.
Also remember that edge services can be hosted in a private cloud as well as public clouds like Amazon web services, google cloud, and Microsoft Azure.
The Service provider Edge provides services over the fixed/mobile network infrastructure. This is a shared infrastructure which means it is not dedicated to one end user. CSPs can leverage their fixed and mobile networks at the edge and provide edge platforms closer to the user.
Access Edge Layer: closest to end device, can be zero or one hop away from the last mile network
Aggregation Edge Layer: The layer which is one hop away from access edge layer
Regional Edge: Often closer to Access Edge than the centralized data center.
User Edge
This is on the other side of the last mile network. Normally this is dedicated and customer-owned. Though it can be service provider owned/managed but dedicated for a single customer
Device Edge : Edge computing capabilities on the device or the user side of the last mile network. Often depends on a gateway in the field to collect and process data from devices. It may have limited compute/storage from user devices( phones, laptops , sensors etc)
Constrained Device Edge : This category includes micro controller based devices which are highly distributed. They can range from simple functions sensors that have no compute to programmable PLCs that have some compute capabilities
Smart Device Edge: This includes IOT gateways, smart phones and PCs.
Together the constrained edge devices and Smart Device Edge represent the “things” in IoT
After discussing Edge location, it is the right time to bring in a discussion of Multi-Access Edge Computing ( MEC)
Difference of Edge Computing and MEC
This diagram from LF Edge ( Linux Foundation) clearly shows, “MEC versus Edge computing”.
Edge computing is an umbrella word that includes MEC as a subset. MEC refers to the Telco Edge or the service provider Edge. ( MEC terminology/standards come from ETSI, initially called “Mobile Edge Computing” later on renamed as “Multi-Access Edge computing)
The end-to-end computing that includes User Edge and Service Provider Edge covers the complete scope of Edge computing.
However, MEC covers a subset scope and refers to the Edge computing provided by a service provider.
As shown here, the MEC cloud is offered by the network operator who owns the network. LF Edge also calls it Telco 5G Edge if the operator uses its 5G network to provide MEC services.
However, do note that MEC extends somewhat in the User Edge area also. This would be the use case where a Telco provides CPE at the customer site.
And do you know that operators have big leverage in the edge computing industry!
They own the “access network” which gives them leverage over the others. They own the 4G/5G network or broadband network for that matter which neither enterprises have nor the hyperscalers have. This is important as MEC had direct access to the network/radio quality, based on which MEC can take intelligent actions.
Neither Hyperscalers like AWS or AZURE have this kind of access network.
So if hyperscalers would like to deploy their Edge platform locally, they need to work out some sort of collaboration with the operators to use their network/physical sites and have a win-win business model for both.It has a huge monetization potential for Telcos
Use Cases of Edge Cloud
I already mentioned the big monetization potential of Edge computing/MEC. It would make sense to discuss the uses cases at this stage.
As you would see that with edge computing it is possible to bring intelligence to the remote locations. These are only few of the many use cases Edge computing has.
Gaming:
Edge computing enables placing gaming servers closer to the users, thus it is possible to reduce latency and provide a fully responsive gaming experience
Autonomous Vehicles/Self-Driving cars:
Autonomous vehicles require Edge computing servers for extremely low latency, so real-time action can be taken to prevent an accident
CDN and Caching
Video is the most popular service on the internet. Users may face low quality of service because of latency and limited bandwidth to reach remote video servers. Bringing video CDN closer to users can improve the quality of experience of the users. This is not limited to video, but any content can be brought to the network edge.
IoT and Big Data
Edge computing can facilitate computation and storage resources for IoT and Big data closer to the user ensuring a fast response to user requests
Industrial IoT ( IIoT)
With edge computing, industrial operators can perform critical analysis closer to sensors and machines reducing the latency for machine decision making.
Telemedicine
Remote medical diagnosis using telemedicine services will become more commonplace if we have better connectivity between patients and doctors. With edge computing, this becomes easier.
Smart Cities Services
In order to make cities smarter, there should be a lot of sensor nodes installed around the city. These sensors collect various kinds of data from different sources such as weather, air pollution, water level, etc..
Augmented Reality (AR)
AR can provide an immersive media experience to the users. For example in sports stadiums, “virtual cameras” present views from within the field of play, giving spectators the experiences from the perspective of the players themselves.
That’s it about an introductory guide on “what is edge computing and MEC”. With an edge computing model, service providers can monetize their networks effectively and improve customer experiences.
How do you view the potential of the technology, let me know in the comments below.
Software-defined Data Center is no longer a “marketing term” as many would make you believe.
Rather, it is a complete framework for building your next software-based data center backed by standards.
So it is important to understand a Software-defined data center ( SDDC)! But that’s not the only reason.
The other reason is the confusion out there in knowing the difference between SDCC and cloud.
Yes !
Many do not know if their applications are in SDCC or in the cloud. Whether they are establishing a cloud or SDCC.
And many still think the “Data Center”, as an enclosed four-walled structure housing communication cabinets within our “access”?
However, the data center has evolved since then. Applications have moved beyond a certain location to cloud somewhere else beyond our access. A Software-Defined data center is NOT just on-premises but has access to public (public cloud services) or works as a hybrid cloud.
So I thought let me comprehensively cover all these topics in a “Software-Defined Data Center tutorial” for you.
For example:
I know a lot of you will be interested especially in the last item in the table of contents, I am keeping it intentionally in the end as you can appreciate the differences once you know exactly what is SDDC. So you if know what is SDDC you can jump straight to the end, else you can follow the sequence of this blog.
What is Software-defined Data Center ?
SDDC started as a marketing term by one of the vendors back in 2012, however, it has taken off considerably after that with one of the standard body DMTF involve to defined the relevant standards related to Open Software-defined data center.
According to DMTF open Software-defined Data center is defined as follows:
“A programmatic abstraction of logical compute, network, storage, and other resources, represented as software. These resources are dynamically discovered, provisioned, and configured based on workload requirements. Thus, the SDDC enables policy-driven orchestration of workloads, as well as measurement and management of resources consumed”
“A data storage facility in which all infrastructure elements—networking, storage, CPU, and security—are virtualized and delivered as a service. Deployment, operation, provisioning, and configuration are abstracted from hardware”
Before we dig deeper into SDDC, it makes sense to know the difference between the traditional data center and SDDC
Difference between Traditional data center and Software-defined data Center
Traditionally data centers use physical infrastructure like physical servers, switches, and storage resources. Their scalability is located individually to each hardware element on site.( servers, switches, firewalls, storage systems, etc). Software-Defined Data Center uses “virtualization” to abstract all these hardware resources on-site providing highly scalable, efficient, and portable virtual compute, networking, networking, and security.
To understand SDDC, understanding Virtualization and Hypervisor is a MUST
With virtualization, we make a software version of something like compute, storage, and networking applications.
What makes virtualization feasible is the “Hypervisor”
The hypervisor is a piece of software that runs on top of a server. It divides the resources of the physical resource and allocates them in the virtual environment. So with the hypervisor, we can turn a physical server into virtual machines ( VMs) with dedicated CPUs, memory, and operating systems.
Once we have the VMs, Instead of having one application on the physical server, we can have multiple applications on the same server resulting in efficiency and cost savings. This is also called virtualization.
Components of Software Defined Data Center
Software Define Compute ( Compute virtualization)
This is the first step towards the SDDC and is also explained under the hypervisor above; this is also called server/compute virtualization or physical hardware virtualization. This lets you run virtual servers on top of a physical server. In simple terms, the CPU and memory of the physical server is allocated to the virtual server
Software Define Network ( Network Virtualization)
The software-Defined networking enables network abstraction and lets you provision and run networks independent of the hardware networking components. One of the challenges with the growth in virtual machines is that the current networks do not facilitate the migration of VMs from one DC to another DC. The IP addresses of VMs are tightly coupled to the physical networks which makes migration very complex. To solve this issue, network virtualization enables virtual overlays that run on top of the physical network/underlay. . This overlay enables hiding of the IP addresses from the physical underlay network thus making the migration of VMs, a breeze. In addition network virtualization brings flexibility and open doors for innovation as new services can be launched without any dependence on the upgrade of the networking hardware
SDN in Data Centers
Software Defined Storage ( Storage Virtualization)
Software-defined storage separates storage software from its hardware. SDS runs on industry-standard x86 servers versus the traditional NAS or SAN systems. Decoupling storage software from hardware enables a lot of flexibility. The storage capacity can be easily expanded as there is a need for expansion.
Storage has come of age. Traditional monolithic storage is sold as a bundle of industry-specific hardware and proprietary software. With the SDS, there is no need for specific hardware, also the SDS adds a software layer between the physical storage and the data request, This enables the use of APIs to manage and maintain the storage of devices. The storage can be scaled out easily while automation can bring the costs down.
Software-Defined Storage
Automation and Orchestration layer
Simply virtualizing functions is not enough. With so many moving pieces in an SDDC, it is mandatory to have a robust automation and orchestration layer. Automation refers to automating a single task like spinning up a VM while orchestration refers to automating a collection of tasks in a certain sequence like spinning up a VM, assigning an IP address then creating a virtual network, etc. A central Orchestration and automation layer can be used to efficiently allocate resources, configure them, update them, monitor operations and take autonomous actions based on close loop controls.
SDDC Architecture:
SDDC architecture as provided by the DTMF is shown in the figure below.
Few of the points related to the architecture
1. At the lower layer is the resources. The resources shown are storage, network, and compute. There can be other software and services ( for example security components like firewalls, IPS, IDS to facilitate security as a service) in addition to the external cloud, which can be a public cloud.
2. One of the most important layers is the “DAL” i.e Datacenter Abstraction layer. The DAL layer abstracts the resources towards the users at the upper layers. This abstraction is done in the standard way providing standard APIs. For example, DTMF has defined the common information models, CMDBf, and OVF formats.
3. The management of the resources is done through SDDC management automation software that has an end-to-end view of the resources. The management interface is defined in CIMI ( Cloud infrastructure Management Interface)
Resource Pooling helps save costs. Instead of buying individual servers and networking hardware, which can over-dimension the hardware, the same hardware can be partitioned using virtualization. Multiple VMs, for example, can be hosted on a single server instead of spinning up a server for each new application.
Scalability and Elasticity
Seamless ability to scale the infrastructure as and when desired. Elastic resources to scale up and scale-out on-demand brings high scale scalability
Agility & Automation
The time to provision services is decreased. It does not take days and months to provision a server, an application, and configure networking. All are software-based which can be done instantaneously. Virtualization combined with automation/orchestration is a real-time saver and opens the door for innovation. The automation layer can be used to efficiently allocate resources, configure them, update them, monitor operations and take autonomous actions based on close loop controls.
APIs & Programmability
Simplified data center management is another benefit. There are common information models through which resources can be programmed facilitating services management through a single dashboard to internal or external parties.
Software-defined Data Center vs Cloud
In order to understand this difference, it is important to refresh the definition of cloud.
What is Cloud?
Let’s take the definition of cloud according to NIST
“cloud computing is a model for enabling ubiquitous, convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, servers, storage, applications, and services) that can be rapidly provisioned and released with minimal management effort or service provider interaction.”
The NIST definition lists five essential characteristics of cloud computing:
on-demand, self-service,
broad network access,
resource pooling,
rapid elasticity or expansion,
measured service.
It also lists three “service models” (software, platform, and infrastructure), and four “deployment models” (private cloud, community, public, and hybrid) that together categorize ways to deliver cloud services”.
Difference with SDCC
By comparing the earlier definitions and architecture of SDDC with that of Cloud it is clear that SDDC focuses more on the architecture and defining the standard interfaces ( for example datacenter abstraction layer) while cloud is more focused on the services and capabilities. However, if you read the definition of cloud clearly, there is nothing in this definition that the SDDC is not able to provide.
Therefore we can say that “Software-defined data center (SDDC) provides the components and architecture to build a cloud. Further, Open SDDC is one way to build the cloud. But there could be other ways to build the cloud, for example, any vendor-specific architecture.
To make it simple we use “SDDC” to build “cloud”
Its your turn to tell me what is your understanding of SDDC and cloud? and do you agree with this way of explanation?