Short answer
A robots.txt file is a plain text file at yourstore.in/robots.txt. It starts with `User-agent: *`, then lists `Disallow:` lines for paths crawlers should skip, then a `Sitemap:` line. Use the free robots.txt generator on this site to build one for your platform and to choose which AI crawlers to block.
Free tools for this
What the file does
I
t
i
s
a
p
o
l
i
t
e
r
e
q
u
e
s
t
,
n
o
t
a
l
o
c
k
.
B
e
f
o
r
e
a
c
r
a
w
l
e
r
r
e
a
d
s
y
o
u
r
p
a
g
e
s
i
t
a
s
k
s
r
o
b
o
t
s
.
t
x
t
w
h
a
t
i
t
m
a
y
v
i
s
i
t
.
S
e
a
r
c
h
e
n
g
i
n
e
s
a
n
d
t
h
e
b
i
g
A
I
c
o
m
p
a
n
i
e
s
f
o
l
l
o
w
i
t
.
A
b
a
d
b
o
t
c
a
n
i
g
n
o
r
e
i
t
,
a
n
d
t
h
e
f
i
l
e
i
s
p
u
b
l
i
c
,
s
o
n
e
v
e
r
l
i
s
t
a
p
r
i
v
a
t
e
a
d
d
r
e
s
s
i
n
i
t
.
The three parts you need
User-agent says who the rules are for. * means everyone. A named bot, such as GPTBot, means just that crawler.
Disallow lists paths to skip, such as /cart or /checkout. Allow makes an exception inside a blocked area, for example WordPress's /wp-admin/admin-ajax.php.
Sitemap points to your sitemap, so crawlers find every page you want found.
Try it now: Robots.txt Generator → · AI Crawler Checker →
What to block on a store
Pages with no search value: cart, checkout, account, internal search results, and filter and sort addresses that make endless near-copies of a collection.
Do not block your product, collection or blog pages, your images, or the CSS and scripts the pages need to display. Google reads those to understand the page.
Blocking AI crawlers
There are three kinds. Training crawlers (GPTBot, ClaudeBot, Google-Extended) collect pages to train models. Search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot) build the index behind AI search. Assistant crawlers (ChatGPT-User, Claude-User, Perplexity-User) open a page when a person asks about it.
Blocking training crawlers does not change your Google rankings. Blocking the search and assistant crawlers keeps you out of those answers. Decide per kind, not all or nothing.
The mistake that hides a whole site
Disallow: / under User-agent: * blocks every crawler from everything. It is meant for a site under construction, and it often survives a launch. Check yours: open yourstore.in/robots.txt and look for a lone slash.
Also remember robots.txt does not remove a page from Google. To keep a page out of search, use a noindex tag and leave the page crawlable so the tag can be read.
On Shopify, WordPress and WooCommerce
Shopify already writes a sensible default. You change it by adding a robots.txt.liquid file to your theme, and you add rules rather than replace the whole file.
WordPress and WooCommerce let you upload a file to the site root, or edit it from your SEO plugin.