使用HtmlAgilityPack抓取網址內容會產生錯誤

asp.net c# html-agility-pack webforms

我正在使用HtmlAgilityPack從網址中獲取文本,該文本適用於大多數網站,對於某些網站來說,它今天開始返回錯誤。

錯誤是以下行代碼doc = webGet.Load(url);錯誤消息: The underlying connection was closed: An unexpected error occurred on a send.

不知道為什麼我得到這個錯誤,因為它正在使用此網站url以前的示例url: link

我試過像bbc.com這樣的https url,它適用於它。任何指針,如果他們的代碼有問題

 HtmlDocument doc = new HtmlDocument();
            var url = txtGrabNewsURL.Text.Trim();

        var webGet = new HtmlWeb();
        doc = webGet.Load(url);
        var baseUrl = new Uri(url);
        //  doc.LoadHtml(response);

        String title = (from x in doc.DocumentNode.Descendants()
                        where x.Name.ToLower() == "title"
                        select x.InnerText).FirstOrDefault();

        String desc = (from x in doc.DocumentNode.Descendants()
                       where x.Name.ToLower() == "meta"
                       && x.Attributes["name"] != null
                       && x.Attributes["name"].Value.ToLower() == "description"
                       select x.Attributes["content"].Value).FirstOrDefault();

        String ogImage = (from x in doc.DocumentNode.Descendants()
                          where x.Name.ToLower() == "meta"
                          && x.Attributes["property"] != null
                          && x.Attributes["property"].Value.ToLower() == "og:image"
                          select x.Attributes["content"].Value).FirstOrDefault();


        List<String> imgs = (from x in doc.DocumentNode.Descendants()
                             where x.Name.ToLower() == "img"
                              && x.Attributes["src"] != null
                             select x.Attributes["src"].Value).ToList<String>();

        List<String> imgList = (from x in doc.DocumentNode.Descendants("img")
                                where x.Attributes["src"] != null
                                select x.Attributes["src"].Value.ToLower()).ToList<String>();

完整錯誤詳情

System.Net.WebException was caught
  HResult=-2146233079
  Message=The underlying connection was closed: An unexpected error occurred on a send.
  Source=System
  StackTrace:
       at System.Net.HttpWebRequest.GetResponse()
       at HtmlAgilityPack.HtmlWeb.Get(Uri uri, String method, String path, HtmlDocument doc, IWebProxy proxy, ICredentials creds) in D:\Source\htmlagilitypack.new\Trunk\HtmlAgilityPack\HtmlWeb.cs:line 1355
       at HtmlAgilityPack.HtmlWeb.LoadUrl(Uri uri, String method, WebProxy proxy, NetworkCredential creds) in D:\Source\htmlagilitypack.new\Trunk\HtmlAgilityPack\HtmlWeb.cs:line 1479
       at HtmlAgilityPack.HtmlWeb.Load(String url, String method) in D:\Source\htmlagilitypack.new\Trunk\HtmlAgilityPack\HtmlWeb.cs:line 1106
       at HtmlAgilityPack.HtmlWeb.Load(String url) in D:\Source\htmlagilitypack.new\Trunk\HtmlAgilityPack\HtmlWeb.cs:line 1061
       at _admin_News.btnGrabNews_Click(Object sender, EventArgs e) in c:\path\News.aspx.cs:line 361
  InnerException: System.IO.IOException
       HResult=-2146232800
       Message=Authentication failed because the remote party has closed the transport stream.
       Source=System
       StackTrace:
            at System.Net.Security.SslState.StartReadFrame(Byte[] buffer, Int32 readBytes, AsyncProtocolRequest asyncRequest)
            at System.Net.Security.SslState.StartReceiveBlob(Byte[] buffer, AsyncProtocolRequest asyncRequest)
            at System.Net.Security.SslState.CheckCompletionBeforeNextReceive(ProtocolToken message, AsyncProtocolRequest asyncRequest)
            at System.Net.Security.SslState.StartSendBlob(Byte[] incoming, Int32 count, AsyncProtocolRequest asyncRequest)
            at System.Net.Security.SslState.ForceAuthentication(Boolean receiveFirst, Byte[] buffer, AsyncProtocolRequest asyncRequest)
            at System.Net.Security.SslState.ProcessAuthentication(LazyAsyncResult lazyResult)
            at System.Net.TlsStream.CallProcessAuthentication(Object state)
            at System.Threading.ExecutionContext.RunInternal(ExecutionContext executionContext, ContextCallback callback, Object state, Boolean preserveSyncCtx)
            at System.Threading.ExecutionContext.Run(ExecutionContext executionContext, ContextCallback callback, Object state, Boolean preserveSyncCtx)
            at System.Threading.ExecutionContext.Run(ExecutionContext executionContext, ContextCallback callback, Object state)
            at System.Net.TlsStream.ProcessAuthentication(LazyAsyncResult result)
            at System.Net.TlsStream.Write(Byte[] buffer, Int32 offset, Int32 size)
            at System.Net.PooledStream.Write(Byte[] buffer, Int32 offset, Int32 size)
            at System.Net.ConnectStream.WriteHeaders(Boolean async)
       InnerException: 

一般承認的答案

如果它只發生在HTTP S資源上,那麼你的目標是.Net 4,那麼它可能與默認的SSL / TLS支持有關。請嘗試以下方法:

using System.Net;

static void Main()
{
    //place this anywhere in your code prior to invoking the Web request
    ServicePointManager.SecurityProtocol = SecurityProtocolType.Tls | SecurityProtocolType.Tls11 | SecurityProtocolType.Tls12 |  SecurityProtocolType.Ssl3; 
}

熱門答案

我在我的本地機器上運行代碼,它工作正常,輸出沒有任何錯誤。我以為時間網站不工作,發生了連接問題。

   HtmlDocument doc = new HtmlDocument();
        var url = "https://m.gulfnews.com/business/sectors/banking/rebuilding-lives-10-years-after-lehman-s-fall-1.2277318"

    var webGet = new HtmlWeb();
    doc = webGet.Load(url);

    String title = (from x in doc.DocumentNode.Descendants()
                    where x.Name.ToLower() == "title"
                    select x.InnerText).FirstOrDefault();

產出:重建生命,10。 。 。 。 。等等。



許可下: CC-BY-SA with attribution
不隸屬於 Stack Overflow
這個KB合法嗎? 是的,了解原因
許可下: CC-BY-SA with attribution
不隸屬於 Stack Overflow
這個KB合法嗎? 是的,了解原因