Bỏ qua để đến Nội dung
Menu
Câu hỏi này đã bị gắn cờ
1 Trả lời
3876 Lượt xem

Hello,


Google can't for some reason fetch generated robots.txt, although it should be able to do so. I can access robots.txt using curl with following output:

curl http://www.mydomain.com/robots.txt

User-agent: *
Disallow: /web/login
Allow: *

User-Agent: Googlebot
Disallow: /web/login


Google is complaining:


Failed: Robots.txt unreachable

Any idea what is wrong?

Also, because of that Google can't access sitemap.xml. 

Another problem is about sitemap.xml. I contains URL's with http, not https prefix. They are valid, as we have http->https redirection rule, but I would prefer to have it correctly in sitemap in the first place. Any help with that?


Many thanks in advance.


Lumir

Ảnh đại diện
Huỷ bỏ
Tác giả Câu trả lời hay nhất

I have found a problem and fixed it. In our case problem was, that we had issue with Nginx proxy settings. When we were accessing the our domain webpages using curl (command line) or Safari, everything seems to be working. But when we tried access website using Firefox, we received an SSL error:

SSL_ERROR_RX_UNEXPECTED_NEW_SESSION_TICKET

We had to move from all sites handled by the Nginx line 

ssl_session_tickets off;

to the /etc/nginx/nginx.conf, section http {} and restart the Nginx.

More info here https://serverfault.com/questions/1021041/browsers-reported-ssl-error-when-one-of-the-server-blocks-in-nginx-configur

This was preventing Google accessing URL's of the domain. Now it's fixed.


Ảnh đại diện
Huỷ bỏ
Bài viết liên quan Trả lời Lượt xem Hoạt động
1
thg 1 23
3502
1
thg 6 17
6019
2
thg 7 15
9079
1
thg 2 25
1004
2
thg 12 24
5941