公式動画ピックアップ
AAPL
ADBE
ADSK
AIG
AMGN
AMZN
BABA
BAC
BL
BOX
C
CHGG
CLDR
COKE
COUP
CRM
CROX
DDOG
DELL
DIS
DOCU
DOMO
ESTC
F
FIVN
GILD
GRUB
GS
GSK
H
HD
HON
HPE
HSBC
IBM
INST
INTC
INTU
IRBT
JCOM
JNJ
JPM
LLY
LMT
M
MA
MCD
MDB
MGM
MMM
MSFT
MSI
NCR
NEM
NEWR
NFLX
NKE
NOW
NTNX
NVDA
NYT
OKTA
ORCL
PD
PG
PLAN
PS
RHT
RNG
SAP
SBUX
SHOP
SMAR
SPLK
SQ
TDOC
TEAM
TSLA
TWOU
TWTR
TXN
UA
UAL
UL
UTX
V
VEEV
VZ
WDAY
WFC
WK
WMT
WORK
YELP
ZEN
ZM
ZS
ZUO
公式動画&関連する動画 [Hosting your own AI models: The hidden cost enterprises miss]
Running your own open weight models to reduce AI token costs sounds like a smart move. GPU time per token is cheaper than what you pay OpenAI, Anthropic, or other hosted model providers. The math seems to work. Until you account for the hours nobody is using them.
In this AI Explainer, Box CTO Ben Kus breaks down why self-hosting open weight models often ends up more expensive for enterprises than using hosted models on demand, and what to do instead.
The core issue is utilization. AI usage is peaky. To avoid making your teams wait during high-demand periods, you need enough GPUs to handle peak load. During off-peak hours, you are paying for GPU capacity that sits idle. That idle time destroys the total cost of ownership calculation that made self-hosting look attractive in the first place.
Hyperscalers and AI labs avoid this problem by spreading utilization across thousands of customers simultaneously. They run their hardware efficiently, apply continuous optimizations, and pass those savings on over time. A single enterprise cannot replicate that economics regardless of how well its infrastructure team executes.
The practical takeaway: for most enterprises, on-demand hosted models, whether accessed directly from a lab or through a platform like Box that procures and manages them on your behalf, will end up cheaper than self-hosting when you account for the full picture, including idle GPU costs, infrastructure management, and the ongoing optimization work that hosted providers do at scale.
This is part of Box's AI Explainer Series: short, practical videos that help enterprise teams understand how to get the most out of AI in their workflows.
FAQs:
Q: Why do enterprises consider self-hosting open weight models in the first place?
A: Open weight models are available without licensing costs, and the raw GPU compute cost per token is lower than what you pay for hosted models from providers like OpenAI or Anthropic. On paper, self-hosting appears to be a cost-saving move.
Q: What is the hidden cost of self-hosting AI models?
A: AI usage is peaky. To avoid delays during high-demand periods, enterprises need enough GPUs to handle peak load. During off-peak hours, those GPUs sit idle but still cost money. That idle time significantly increases the total cost of ownership and often makes self-hosting more expensive than using hosted models on demand.
Q: How do hyperscalers and AI labs keep their costs lower?
A: Providers like OpenAI, Anthropic, and Google spread utilization across thousands of customers simultaneously. That allows them to run their hardware at much higher efficiency rates than any single enterprise could achieve, apply continuous optimizations, and lower costs over time.
Q: When does self-hosting make sense for enterprises?
A: Self-hosting can make sense in specific scenarios, such as when regulatory requirements prevent data from leaving a controlled environment, or when usage is consistently high enough to keep GPU utilization rates competitive with hosted providers. For most enterprises, however, on-demand hosted models are the more cost-effective choice.
Q: How does Box help enterprises manage AI model costs?
A: Box procures and manages access to leading AI models on behalf of enterprise customers, allowing teams to use the right model for the right workflow without building and maintaining their own infrastructure. This gives enterprises the flexibility of model choice without the overhead of self-hosting.
120794
1